Training sample determination method, model training method, electronic equipment and medium
By randomly obtaining and expanding basic text in the text database and determining target instructions, the problem of poor training samples in the prior art is solved, efficient and automatic training sample generation is achieved, and the training efficiency and quality of the instruction generation model is improved.
Patent Information
- Application Number
- CN202510256622.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, when training neural network models through the network through the network, the quality of the training samples obtained is poor, and the efficiency of manual annotation data is low. The quality of the training samples generated by the general large language model is poor and requires manual verification.
Randomly obtain basic text in the text database, expand the basic text to obtain target text, determine the target instructions corresponding to the target text, and determine the training samples based on the target text and the target instructions. This method improves the quality and data volume of training samples by extending the basic text and determining the target instructions.
It realizes automatic and efficient acquisition of high-quality training samples, improves the training efficiency and domain capabilities of instruction generation models, and reduces the burden of manual annotation and verification.
Smart Images

Figure CN120104786A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular to a method for determining a training sample, a model training method, an electronic device, and a medium. Background Art
[0002] At present, natural language can be used to control robots to perform tasks. For example, if a user gives a natural language instruction "put the fruit in the basket", the robot can understand the natural language instruction and execute it correctly. The robot needs to understand the natural language instruction through a neural network model, and the neural network model needs to generate control instructions that the robot can understand from the natural language instruction, and then control the robot to perform tasks through the control instruction.
[0003] In the prior art, the neural network model is trained by acquiring data on the network as training samples. However, the training samples obtained in this way have the problem of poor quality. Summary of the invention
[0004] Various aspects of the present disclosure provide a method for determining a training sample, a model training method, an electronic device, and a medium to solve the technical problem that the training samples obtained in the prior art have poor quality.
[0005] A first aspect of an embodiment of the present disclosure provides a method for determining a training sample, comprising: randomly acquiring a basic text in a text database, the basic text comprising a first placeholder, the semantics of the basic text being to perform a corresponding action on the first placeholder, the text database comprising: a plurality of candidate texts, different candidate texts corresponding to different actions, the basic text being one of the plurality of candidate texts;
[0006] The basic text is expanded to obtain a first target text, the first target text includes: a first extended text obtained by expanding the first placeholder, and the semantics of the first target text is to perform an action on the first extended text;
[0007] Determine a first target instruction of the first target text, the first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to perform an action on the object indicated by the first extended text;
[0008] A training sample is determined according to the first target text and the first target instruction.
[0009] A second aspect of the present disclosure provides a model training method, including:
[0010] Acquire a training sample, where the training sample is obtained according to the method for determining the training sample of the first aspect;
[0011] The instruction generation model is trained according to the training samples to obtain a trained instruction generation model, and the instruction generation model is used to generate instructions for controlling the robot.
[0012] A third aspect of the embodiments of the present disclosure provides a device for determining a training sample, including:
[0013] An acquisition module is used to randomly acquire a basic text from a text database, the basic text includes a first placeholder, the semantics of the basic text is to execute a corresponding action for the first placeholder, the text database includes: a plurality of candidate texts, different candidate texts correspond to different actions, and the basic text is one of the plurality of candidate texts;
[0014] An expansion module, configured to expand a basic text to obtain a first target text, wherein the first target text includes: a first extended text obtained by expanding a first placeholder, and the semantics of the first target text is to perform an action on the first extended text;
[0015] A first determination module is used to determine a first target instruction of a first target text, the first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to perform an action on an object indicated by the first extended text;
[0016] The second determination module is used to determine a training sample according to the first target text and the first target instruction.
[0017] A fourth aspect of the present disclosure provides a model training device, including:
[0018] An acquisition module, used to acquire a training sample, where the training sample is obtained according to the method for determining the training sample of the first aspect;
[0019] The generation module is used to train the instruction generation model according to the training samples to obtain the trained instruction generation model, and the instruction generation model is used to generate instructions for controlling the robot.
[0020] A fifth aspect of an embodiment of the present disclosure provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the first aspect and / or the second aspect when executing the computer program.
[0021] A sixth aspect of an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method of the first aspect and / or the second aspect.
[0022] A seventh aspect of an embodiment of the present disclosure provides a computer program product, the program product comprising: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the method of the first aspect and / or the second aspect.
[0023] The embodiments of the present disclosure are applied in the scenario of determining training samples. The present disclosure randomly obtains a basic text from a text database, wherein the basic text includes a first placeholder, and the semantics of the basic text is to execute a corresponding action for the first placeholder. The text database includes: multiple candidate texts, different candidate texts correspond to different actions, and the basic text is one of the multiple candidate texts; the basic text is extended to obtain a first target text, and the first target text includes: a first extended text obtained by extending the first placeholder, and the semantics of the first target text is to execute an action for the first extended text; a first target instruction of the first target text is determined, the first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to execute an action for an object indicated by the first extended text; the training sample is determined according to the first target text and the first target instruction, so that a training sample with high quality and a large amount of data can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure. In the drawings:
[0025] Figure 1 An application scenario diagram of a method for determining a training sample provided by an exemplary embodiment of the present disclosure;
[0026] Figure 2 A flowchart of a method for determining a training sample provided by an exemplary embodiment of the present disclosure;
[0027] Figure 3 A flowchart of another method for determining a training sample provided by an exemplary embodiment of the present disclosure;
[0028] Figure 4 A flowchart of a model training method provided for an exemplary embodiment of the present disclosure;
[0029] Figure 5 A structural block diagram of a device for determining a training sample provided by an exemplary embodiment of the present disclosure;
[0030] Figure 6 A structural block diagram of a model training device provided for an exemplary embodiment of the present disclosure;
[0031] Figure 7 A schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below in combination with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0033] In the related art, the following methods are used to determine training samples: (1) Obtaining data from the Internet and processing the data to obtain training samples. However, since the data on the Internet is relatively complex, the quality of the obtained training samples is poor. (2) Using manually annotated data to obtain training samples, wherein manually annotated data has the problem of low efficiency. (3) Using a general LLM (Large Language Model) to generate training samples, wherein, since the general LLM has poor domain-specific capabilities, the quality of the generated training samples is poor, and the generated training samples need to be manually verified, which also has the problem of low efficiency.
[0034] Based on the above problems, the present invention improves a method for determining a training sample, which can automatically obtain a training sample, and the training sample is determined efficiently. The obtained training sample can be applied to the training of the instruction generation model in the field of robots, thereby improving the field capability. Furthermore, by expanding the basic text to obtain a first target text, determining a first target instruction of the first target text, and determining a training sample based on the first target text and the first target instruction, the quality of the training sample can be improved.
[0035] In addition, an application scenario of the embodiment of the present disclosure is as follows Figure 1 , wherein the obtained training samples are used to train the instruction generation model, and then the natural language is input into the trained instruction generation model for processing to obtain the instructions output by the instruction generation model, which can enable the air robot 11 to perform the corresponding task.
[0036] in, Figure 1 This is only an exemplary application scenario, and the embodiment of the present disclosure can be applied to the determination of any training sample. The embodiment of the present disclosure does not limit the specific application scenario.
[0037] Figure 2 A flowchart of a method for determining a training sample provided by an exemplary embodiment of the present disclosure specifically includes the following steps:
[0038] S201, randomly obtaining basic text from a text database.
[0039] The basic text includes a first placeholder, and the semantics of the basic text is to execute a corresponding action for the first placeholder. The text database includes: multiple alternative texts, different alternative texts correspond to different actions, and the basic text is one of the multiple alternative texts.
[0040] In the present disclosure, the candidate texts in the text database are preset, and the actions corresponding to the candidate texts are actions that can be executed by the robot. For example, the candidate texts include: move A to B, put A to B, pick up A, put down A, count A, and locate A. Among them, the action corresponding to moving A to B is "move", the action corresponding to putting A to B is "pick up and put down", the action corresponding to picking up A is "pick up", the action corresponding to putting down A is "put down", the action corresponding to counting A is "counting", and the action corresponding to locating A is "locating".
[0041] Furthermore, the first placeholder may be represented by letters, such as A in the alternative text. In addition, the first placeholder may also be represented by other methods, such as spaces or punctuation marks, which are not limited here.
[0042] In the present disclosure, basic texts are randomly acquired to improve the coverage of subsequently generated training samples and to improve the robustness of the instruction generation model trained by the training samples.
[0043] S202: Expand the basic text to obtain a first target text.
[0044] The first target text includes: a first extended text obtained by expanding the first placeholder, and the semantics of the first target text is to execute an action on the first extended text.
[0045] In the present disclosure, the first target text is expanded after the basic text, and the action corresponding to the first target text is the same as the basic text. For example, the basic text is: put A to B, and the action corresponding to the basic text is "pick up, put down", then the action corresponding to the first target text is also "pick up, put down". The first target text is such as: put 2 As to B, or put A of C to B.
[0046] It can be understood that the purpose of extending the basic text is to increase the richness of the training samples, wherein the actions corresponding to the basic text after expansion are still actions that can be executed by the robot.
[0047] In the present disclosure, one way to expand the basic text is to expand only the first placeholder in the basic text, and retain the other characters in the basic text except the placeholder in the first target text. For example, the basic text is to put A into B, and the first placeholder is "A". The first expanded text obtained by expanding the first placeholder is A of C, 2 As, A to the left of C, or A inside C. It can be understood that the first expanded text can be obtained by expanding the first placeholder in any form, which is not limited here.
[0048] S203: Determine a first target instruction of a first target text.
[0049] The first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to perform an action on the object indicated by the first extended text.
[0050] Among them, the format of the first target instruction is preset and is an instruction that can be executed by the robot.
[0051] For example, if the first target text is to put 2 A's into B, then the corresponding first target instruction is: pick_up(2 A's); place_to(B).
[0052] S204: Determine a training sample according to the first target text and the first target instruction.
[0053] In one embodiment, the training sample includes: a first target text and a first target instruction. In the process of training the instruction generation model, the first target text is input into the instruction generation model to obtain a predicted instruction. The first target instruction is used as label data of the first target text. The loss values of the first target instruction and the predicted instruction are determined, and the model parameters of the instruction generation model are adjusted according to the loss values.
[0054] In the embodiment of the present disclosure, S201 to S204 are executed in a loop, thereby generating multiple training samples, thereby improving the generation efficiency and quality of the training samples.
[0055] In summary, the present disclosure can achieve automatic acquisition of training samples, high efficiency in determining training samples, and the obtained training samples can be applied to the training of instruction generation models in the field of robots, thereby improving the field capability. Furthermore, by expanding the basic text to obtain the first target text, determining the first target instruction of the first target text, and determining the training samples based on the first target text and the first target instruction, the quality of the training samples can be improved.
[0056] Figure 3 A flowchart of another method for determining a training sample provided by an exemplary embodiment of the present disclosure specifically includes the following steps:
[0057] S301, randomly obtaining basic text from a text database.
[0058] The specific implementation process of this step can be referred to S201 and will not be repeated here.
[0059] S302: Expand the first placeholder to obtain a first extended text.
[0060] The first extended text includes: at least one second placeholder. The second placeholder can also be represented by letters. For example, if the first placeholder is "A", the first extended text is "2 A's", the second placeholder is also "A", or if the first extended text is "C's A", the first extended text includes: two second placeholders, namely "C" and "A".
[0061] Specifically, expanding the first placeholder to obtain the first extended text includes: randomly obtaining a first extended rule in an extended rule library, wherein the extended rule library includes: multiple second extended rules, the semantics of the second extended rule is to expand the first placeholder into the corresponding second extended text, different second extended rules correspond to different second extended texts, and the first extended rule is one of the second extended rules; determining that the second extended text corresponding to the first extended rule is the first extended text.
[0062] In the embodiment of the present disclosure, the second expansion rule included in the expansion rule base is preset. For example, the multiple second expansion rules included in the expansion rule base include: A is expanded to A1 and A2, A is expanded to 2 As, A is expanded to A of C, and A is expanded to A in C.
[0063] In addition, the rule extension library may also include: non-extension rules. If S302 is executed and the non-extension rules are randomly selected, the basic text may be subsequently determined as the first target text for generating training samples.
[0064] For example, the first placeholder is A. If A is randomly expanded into A1 and A2 in the expansion rule base, then A1 and A2 can be determined as the first expansion text.
[0065] In the present disclosure, a first random extension rule is added to the extension rule base, which can increase the richness of training samples.
[0066] S303: Use the first extended text to replace the first placeholder in the basic text to obtain a first target text.
[0067] In the disclosed embodiment, the first extended text is used to directly replace the first placeholder in the basic text to obtain the first target text.
[0068] For example, the basic text is "put A into B", the first placeholder is "A", and the first randomly obtained expansion rule is "A expands to C's A", then the first expanded text is determined to be "C's A", and "C's A" is used to replace the "A" in "put A into B", and the first target text obtained is "put C's A into B".
[0069] S304, obtaining a first instruction of a basic text.
[0070] The first instruction includes a first placeholder, and the semantics of the first instruction is the same as that of the basic text.
[0071] In the disclosed embodiment, the first instruction of the basic text can be pre-written into a text database. A correspondence between each alternative text and the first instruction is established in the text database, and then the first instruction of the basic text is determined based on the correspondence. For example, the alternative text is "put A to B", and the corresponding first instruction is "pick_up(A); place_to(B)" or "pick up(A); put down(B)". The alternative text is "pick up A", and the corresponding first instruction is "pick_up(A)" or "pick up(A)". The alternative text is "put down A", and the corresponding first instruction is "place_to(A)" or "put down(A)". The alternative text is "count A", and the corresponding first instruction is "count(A)" or "count(A)". The alternative text is "locate A", and the corresponding first instruction is "locate(A)" or "locate(A)".
[0072] S305: Use the first extended text to replace the first placeholder in the first instruction to obtain a first target instruction.
[0073] In the present disclosure, the first extended text is used to directly replace the first placeholder in the first instruction to obtain the first target text.
[0074] For example, the first instruction is "pick_up(A); place_to(B)", the first extended text is "C's A", and "C's A" is used to replace the "A" in "pick_up(A); place_to(B)", and the first target instruction obtained is "pick_up(C's A); place_to(B)".
[0075] S306: Combine at least two first target texts to obtain a second target text.
[0076] In the embodiment of the present disclosure, S301 to S305 may be executed in a loop to obtain at least two first target texts.
[0077] For example, the first target text P1 is "put A of C to B", and the first target text P2 is "pick up 2 As", then the second target text "put A of C to B, pick up 2 As" is obtained by combining the first target text P1 and the first target text P2.
[0078] Furthermore, if the second placeholders in different first target texts are represented by the same letters, other letters can be used to replace the same letters, so that the second placeholders in different first target texts are represented by different letters. For example, if the second placeholder "A" included in the first target text P1 is the same as the second placeholder "A" included in the first target text P2, the second placeholder "A" included in the first target text P1 can be replaced with "A1", and the second placeholder "A" included in the first target text P2 can be replaced with "A2". Further, the second target text obtained is "Put A1 of C to B, and pick up 2 A2s".
[0079] In the disclosed embodiment, each time S301 is executed, the basic text is randomly obtained, and the first extended text is also random. Therefore, the randomness of the second target text can be improved. Further, by executing S301 to S308 multiple times, different training samples can be obtained, thereby improving the richness of the training samples.
[0080] S307, combining the first target instructions corresponding to at least two first target texts to determine a second target instruction.
[0081] In the embodiment of the present disclosure, S301 to S305 may be executed in a loop to obtain a first target instruction corresponding to each first target text in at least two first target texts.
[0082] For example, the first target instruction corresponding to the first target text P1 "put C's A to B" is "pick_up(C's A); place_to(B)", and the first target instruction corresponding to the first target text P2 "pick up 2 As" is "pick_up(2 A's)", then the second target instruction obtained by combining these two first target instructions is "pick_up(C's A); place_to(B); pick_up(2 A's)".
[0083] Further, referring to the determination process of the second target text, the corresponding second target instruction may be “pick_up(C's A1); place_to(B); pick_up(2 A2)”, wherein the second placeholder included in the second target instruction is the same as the second placeholder included in the second target text.
[0084] S308: Determine a training sample according to the second target text and the second target instruction.
[0085] In one embodiment, the training sample includes: a second target text and a second target instruction. During the process of training the instruction generation model, the second target text is input into the instruction generation model to obtain a predicted instruction. The second target instruction is used as label data of the second target text. The loss value of the second target instruction and the predicted instruction is determined, and the model parameters of the instruction generation model are adjusted according to the loss value.
[0086] In another embodiment, a training sample is determined based on a second target text and a second target instruction, including: obtaining at least one first object text; replacing a corresponding second placeholder in the second target text with at least one first object text to obtain a third target text; replacing a second placeholder in the second target instruction with the first object text to obtain a third target instruction; and determining a training sample based on the third target text and the third target instruction.
[0087] In the present disclosure, the first object text represents a specific object. In the present disclosure, an object text library can be preset, wherein the object text library includes object texts of multiple categories, such as common objects: apple, banana, pear, pencil, eraser, ruler, mobile phone, book, pen; common categories: fruit, vegetable, food and stationery.
[0088] Furthermore, the objects included in the object text library can be set according to the specific application scenario of the robot. If the robot is used to process fruits, the object texts included in the object text library include apples, bananas, pears, etc. and fruits. If the robot is used to process stationery, the object texts included in the object text library include pencils, erasers, rulers, books, and pens. In the present disclosure, the specific objects in the object text library are not limited.
[0089] In one embodiment, the number of the acquired first object texts and the number of the second placeholders in the second target text may be the same, and then one of the second placeholders may be replaced by one of the first object texts.
[0090] For example, the second target text is "Put A1 of C to B, pick up 2 A2s", and the second target instruction is "pick_up(A1 of C); place_to(B); pick_up(2 A2s)". The second target text includes three second placeholders, namely: C, A1 and A2, then three first object texts can be randomly obtained, such as ruler, apple and eraser, and then the first object texts are used to replace the second placeholders in the second target text and the second target instruction, and the third target text obtained is "Put the apple of the ruler to B, pick up 2 erasers", and the third target instruction is "pick_up(apple of the ruler); place_to(B); pick_up(2 erasers)".
[0091] It can be understood that in the present disclosure, the semantics of the third target text and the third target instruction are the same.
[0092] Furthermore, it can be determined that the training samples include: a third target text and a third target instruction. In the process of training the instruction generation model, the third target text is input into the instruction generation model to obtain a predicted instruction. The third target instruction is used as label data of the third target text. The loss values of the third target instruction and the predicted instruction are determined, and the model parameters of the instruction generation model are adjusted according to the loss values.
[0093] In addition, obtaining at least one first object text includes: randomly obtaining at least one first object text in an object text library; or, inputting the second target text into a language model for processing to obtain the first object text corresponding to each second placeholder.
[0094] One way is that the first object text is randomly obtained from the object text library, wherein randomly obtaining the first object text can improve the richness of the training samples.
[0095] Another way is to obtain a prompt word, input the prompt word and the second target text into the language model for processing, and obtain the first object text. The prompt word may be "output the first object text according to the input text content, and the first object text needs to be an object name that conforms to the semantics of the text content" or the prompt word may be "output the first object text according to the input text content, and the first object text needs to be an object name that does not conform to the semantics of the text content". In the present disclosure, the prompt word can be set as needed to prompt the language model to output the first object text corresponding to the second target text according to the prompt word.
[0096] In the present disclosure, the first object text is output according to the language model, which can improve the richness of the first object text.
[0097] In the present disclosure, the basic text also includes: a third placeholder, the third placeholder represents the target position corresponding to the action, and after replacing the corresponding second placeholder in the second target text with at least one first object text, it also includes: obtaining at least one second object text; replacing the corresponding second placeholder in the second target text with at least one first object text, and replacing the corresponding third placeholder in the second target text with at least one second object text to obtain a third target text.
[0098] The method for obtaining the second object text may refer to the first object text, which will not be described in detail here.
[0099] It can be understood that if the basic text also includes a third placeholder, the second object text can be obtained to replace the third placeholder to obtain the third target text. For example, the second target text is "Put A1 of C to B, pick up 2 A2s", the third placeholder is "B", and the second object text is "mobile phone". Then the first object text is used to replace the second placeholder. After the second object text replaces the third placeholder, the third target text obtained is "Put the apple of the ruler to the mobile phone, pick up 2 erasers"
[0100] Among them, determining the training sample according to the third target text and the third target instruction includes: processing the third target text through a language model to obtain a natural language text, the natural language text and the third target text represent the same semantics; determining the training sample, the training sample includes the natural language text and the third target instruction, and the third target instruction is the label of the natural language text.
[0101] It can be understood that the language model can polish the third target text, and the obtained natural language text can be a text that is more in line with the user's language. For example, the third target text is "Put the apple on the ruler on the mobile phone, and pick up two erasers." The natural language text obtained after being processed by the language model is "Please put the apple on the ruler on the mobile phone, and then pick up two erasers."
[0102] In the embodiment of the present disclosure, a third object text can also be obtained to replace the second placeholder in the first target text and the first target instruction to obtain a fourth target text, and the fourth object text is obtained to replace the third placeholder in the first target text and the first target instruction to obtain a fourth target instruction. The training sample then includes the fourth target text and the fourth target instruction.
[0103] In summary, the present disclosure can obtain a large number of relatively rich training samples by randomly obtaining basic text, randomly obtaining the first extension rule, and randomly obtaining the first object text and the second object text, after executing S301 to S308 multiple times in a loop, and because the present disclosure does not require manual annotation, the efficiency of obtaining training samples can be improved. In addition, the present disclosure uses the above method to determine the training samples, which can improve the quality of the obtained training samples.
[0104] Figure 4 A flowchart of a model training method provided by an exemplary embodiment of the present disclosure specifically includes the following steps:
[0105] S401, obtaining training samples.
[0106] In the present disclosure, there are multiple training samples, and each training sample can be used to train the instruction generation model.
[0107] A training sample includes: a first target text and a first target instruction, wherein the first target instruction is a label of the first target text.
[0108] Another training sample includes: a second target text and a second target instruction, where the second target instruction is a label of the second target text.
[0109] Another training sample includes: a third target text and a third target instruction, where the third target instruction is a label of the third target text.
[0110] Another type of training sample includes: a fourth target text and a fourth target instruction, where the fourth target instruction is a label of the fourth target text.
[0111] In the present disclosure, the method for determining each training sample can refer to the above content and will not be repeated here.
[0112] S402, training the instruction generation model according to the training samples to obtain a trained instruction generation model.
[0113] In the present disclosure, the instruction generation model may be an LLM, the target text of the training sample is input into the instruction generation model for processing, a predicted instruction is obtained, the loss value of the target instruction and the predicted instruction is determined, and the model parameters of the instruction generation model are adjusted according to the loss value.
[0114] In addition, the model may be generated according to the training sample training instructions in other ways, which are not limited here.
[0115] It can be understood that the present disclosure uses high-quality training samples to train the instruction generation model, so as to improve the quality of the instruction generation model. The trained instruction generation model can generate instructions for controlling the robot according to the natural language input by the user, thereby realizing accurate control of the robot.
[0116] Reference Figure 5 , is a structural block diagram of a training sample determination device 50 provided in the present disclosure, and the training sample determination device 50 includes:
[0117] An acquisition module 501 is used to randomly acquire a basic text from a text database, the basic text includes a first placeholder, the semantics of the basic text is to execute a corresponding action for the first placeholder, the text database includes: a plurality of candidate texts, different candidate texts correspond to different actions, and the basic text is one of the plurality of candidate texts;
[0118] An expansion module 502 is used to expand the basic text to obtain a first target text, the first target text includes: a first extended text obtained by expanding the first placeholder, and the semantics of the first target text is to perform an action on the first extended text;
[0119] A first determination module 503 is used to determine a first target instruction of a first target text, the first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to perform an action on an object indicated by the first extended text;
[0120] The second determination module 504 is used to determine a training sample according to the first target text and the first target instruction.
[0121] In an optional embodiment, the expansion module 502 is specifically configured to expand the first placeholder to obtain a first extended text, wherein the first extended text includes: at least one second placeholder;
[0122] The first extended text is used to replace the first placeholder in the basic text to obtain a first target text.
[0123] In an optional embodiment, when the expansion module 502 expands the first placeholder to obtain the first extended text, it is specifically configured to:
[0124] A first expansion rule is randomly obtained from an expansion rule library, wherein the expansion rule library includes: a plurality of second expansion rules, the semantics of the second expansion rule is to expand the first placeholder into a corresponding second expansion text, different second expansion rules correspond to different second expansion texts, and the first expansion rule is one of the second expansion rules;
[0125] The second extended text corresponding to the first extended rule is determined to be the first extended text.
[0126] In an optional embodiment, the second determination module 504 is specifically configured to combine at least two first target texts to obtain a second target text;
[0127] Combining the first target instructions corresponding to at least two first target texts to determine a second target instruction;
[0128] A training sample is determined according to the second target text and the second target instruction.
[0129] In an optional embodiment, when the second determination module 504 determines the training sample according to the second target text and the second target instruction, it is specifically used to:
[0130] Obtain at least one first object text;
[0131] Using at least one first object text to replace a corresponding second placeholder in the second target text to obtain a third target text;
[0132] Replacing the second placeholder in the second target instruction with the first object text to obtain a third target instruction;
[0133] A training sample is determined according to the third target text and the third target instruction.
[0134] In an optional embodiment, when acquiring at least one first object text, the second determining module 504 is specifically configured to:
[0135] randomly acquiring at least one first object text in an object text library;
[0136] Alternatively, the second target text is input into a language model for processing to obtain a first object text corresponding to each second placeholder.
[0137] In an optional embodiment, when determining the training sample according to the third target text and the third target instruction, the second determination module 504 is specifically configured to:
[0138] The third target text is processed by a language model to obtain a natural language text, wherein the natural language text and the third target text express the same semantics;
[0139] A training sample is determined, where the training sample includes a natural language text and a third target instruction, where the third target instruction is a label of the natural language text.
[0140] In an optional embodiment, the basic text further includes: a third placeholder, the third placeholder indicates a target position corresponding to the action, and after the second determination module 504 replaces the corresponding second placeholder in the second target text with at least one first object text, it is further specifically configured to:
[0141] Obtain at least one second object text;
[0142] At least one first object text is used to replace a corresponding second placeholder in the second target text, and at least one second object text is used to replace a corresponding third placeholder in the second target text to obtain a third target text.
[0143] In an optional embodiment, the first determining module 503 is specifically configured to:
[0144] Obtaining a first instruction of the basic text, the first instruction including a first placeholder, and the semantics of the first instruction being the same as that of the basic text;
[0145] A first extended text is used to replace a first placeholder in a first instruction to obtain a first target instruction.
[0146] The training sample determination device provided in the present disclosure can implement the above-mentioned training sample determination method. Please refer to the above for details and will not be repeated here.
[0147] Reference Figure 6 , is a structural block diagram of a model training device 60 provided by the present disclosure, and the model training device 60 includes:
[0148] An acquisition module 601 is used to acquire a training sample, where the training sample is obtained according to the method for determining the training sample of the first aspect;
[0149] The generation module 602 is used to train the instruction generation model according to the training samples to obtain a trained instruction generation model, and the instruction generation model is used to generate instructions for controlling the robot.
[0150] The model training device provided in the present disclosure can implement the above-mentioned model training method. Please refer to the above for details and will not be repeated here.
[0151] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel, and are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "second", "first", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit "second" and "first" to be different types.
[0152] Figure 7 A schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present disclosure. Figure 7 As shown, the electronic device 70 includes: a processor 71, and a memory 72 communicatively connected to the processor 71, and the memory 72 stores computer-executable instructions.
[0153] Among them, the processor executes the computer execution instructions stored in the memory to implement the training sample determination method and / or defect detection method provided by any of the above method embodiments, and the specific functions and technical effects that can be achieved are not repeated here.
[0154] The embodiment of the present disclosure further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement any of the above methods.
[0155] The embodiments of the present disclosure also provide a computer program product, which includes: a computer program, which is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes any of the above methods.
[0156] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0157] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0159] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to perform some steps of the methods of various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0160] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0161] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0162] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for determining a training sample, characterized in that: include: A basic text is randomly obtained from a text database, wherein the basic text includes a first placeholder, and the semantics of the basic text is to perform a corresponding action on the first placeholder, wherein the text database includes: a plurality of candidate texts, and different candidate texts correspond to different actions, and the basic text is one of the plurality of candidate texts; Expanding the basic text to obtain a first target text, the first target text comprising: a first extended text obtained by expanding the first placeholder, the semantics of the first target text being to perform the action on the first extended text; Determine a first target instruction of the first target text, wherein the first target text and the first target instruction have the same semantics, and the first target instruction is used to control the robot to perform the action on the object indicated by the first extended text; A training sample is determined according to the first target text and the first target instruction.
2. The method according to claim 1, characterized in that: The step of expanding the basic text to obtain a first target text includes: Expanding the first placeholder to obtain the first extended text, wherein the first extended text includes: at least one second placeholder; The first extended text is used to replace the first placeholder in the basic text to obtain a first target text.
3. The method according to claim 2, characterized in that The step of expanding the first placeholder to obtain a first expanded text includes: The first expansion rule is randomly obtained from an expansion rule library, wherein the expansion rule library includes: a plurality of second expansion rules, the semantics of the second expansion rule is to expand the first placeholder into a corresponding second expansion text, different second expansion rules correspond to different second expansion texts, and the first expansion rule is one of the second expansion rules; It is determined that the second extended text corresponding to the first extended rule is the first extended text.
4. The method according to claim 2, characterized in that: The determining of a training sample according to the first target text and the first target instruction includes: combining at least two first target texts to obtain a second target text; Combining the first target instructions corresponding to the at least two first target texts to determine a second target instruction; The training sample is determined according to the second target text and the second target instruction.
5. The method according to claim 4, characterized in that The step of determining the training sample according to the second target text and the second target instruction includes: Obtain at least one first object text; Replacing a corresponding second placeholder in the second target text with the at least one first object text to obtain a third target text; Replacing the second placeholder in the second target instruction with the first object text to obtain a third target instruction; The training sample is determined according to the third target text and the third target instruction.
6. The method according to claim 5, characterized in that The obtaining of at least one first object text comprises: randomly acquiring at least one first object text in an object text library; Alternatively, the second target text is input into a language model for processing to obtain a first object text corresponding to each second placeholder.
7. The method according to claim 5 or 6, characterized in that: The determining the training sample according to the third target text and the third target instruction comprises: Processing the third target text by using the language model to obtain a natural language text, wherein the natural language text and the third target text represent the same semantics; The training sample is determined, where the training sample includes the natural language text and the third target instruction, and the third target instruction is a label of the natural language text.
8. The method according to claim 5 or 6, characterized in that: The basic text further includes: a third placeholder, the third placeholder indicating a target position corresponding to the action, and after the at least one first object text is used to replace the corresponding second placeholder in the second target text, the method further includes: Obtain at least one second object text; The at least one first object text is used to replace the corresponding second placeholder in the second target text, and the at least one second object text is used to replace the corresponding third placeholder in the second target text to obtain the third target text.
9. The method according to any one of claims 1 to 6, characterized in that: The determining of the first target instruction of the first target text comprises: Obtaining a first instruction of the basic text, the first instruction comprising a first placeholder, and the semantics of the first instruction being the same as that of the basic text; The first extended text is used to replace the first placeholder in the first instruction to obtain the first target instruction.
10. A model training method, characterized in that: include: Acquire a training sample, where the training sample is obtained according to the method for determining a training sample according to claims 1 to 9; The instruction generation model is trained according to the training samples to obtain a trained instruction generation model, and the instruction generation model is used to generate instructions for controlling the robot.
11. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 10 when executed by a processor.