A natural language trajectory instruction generation method and device and a storage medium
Patent Information
- Application Number
- CN202310411770.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-17
AI Technical Summary
然而,以前的轨迹指令生成模型往往忽略了小规模数据集对轨迹指令生成模型自身的限制,导致轨迹指令生成模型性能较低,存在生成大量错误伪标签的情况
[0031](1)本发明提出了基于双塔结构的轨迹-指令匹配器,通过自适应学习的方式来对伪数据对的匹配度进行打分,若匹配值小于阈值则将其过滤,保证了伪标签的质量,有效降低了数据噪声与潜在的干扰模型学习的风险,提高了自然语言指令生成的准确性。
Smart Images

Figure CN116522899B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual language multimodal fusion technology, and in particular to a method, apparatus and storage medium for generating natural language trajectory instructions based on semi-supervised learning. Background Technology
[0002] Natural language trajectory instruction generation (NLP) aims to output corresponding instruction descriptions based on observed features of a machine's movement path, and is an important technology for human-computer interaction and autonomous machine exploration. This method is also a crucial data augmentation technique for visual language navigation tasks. Automatically generating instructions based on a NLP trajectory instruction generation model can effectively address the problem of low model generalization caused by small labeled datasets. However, previous trajectory instruction generation models often ignored the limitations imposed by small datasets, resulting in low performance and the generation of numerous erroneous pseudo-labels. Furthermore, the lack of methods to verify the quality of pseudo-labels leads to erroneous trajectory instructions that negatively impact model learning and actual task completion. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, and storage medium for generating natural language trajectory instructions based on semi-supervised learning, thereby improving the generalization and accuracy of the model.
[0004] The objective of this invention can be achieved through the following technical solutions:
[0005] A method for generating natural language trajectory instructions based on semi-supervised learning includes the following steps:
[0006] Step 1) Construct a trajectory-instruction generator based on an encoder-decoder structure;
[0007] Step 2) Construct a trajectory-instruction matcher based on a dual-tower structure;
[0008] Step 3) Collect several candidate navigation points in the environment, generate a limited number of trajectories and label them with corresponding natural language instructions. Manually label the data to form a labeled dataset. At the same time, randomly generate a large number of trajectory routes to form an unlabeled dataset.
[0009] Step 4) Train the trajectory-instruction generator and the trajectory-instruction matcher using the labeled dataset;
[0010] Step 5) Based on the unlabeled dataset, generate corresponding pseudo-labels using the trajectory-instruction generator, and filter out low-quality pseudo-labels using the trajectory-instruction matcher.
[0011] Step 6) Merge the filtered high-quality pseudo-label dataset with the labeled dataset to refine the trajectory-instruction generator and obtain a high-performance trajectory-instruction generator.
[0012] Step 7) Repeat steps 5) and 6) until the trajectory-instruction matcher determines that there are no low-quality pseudo-labels, or the maximum number of repetition rounds is reached.
[0013] Furthermore, the trajectory-instruction generator takes as input a set of environmental visual images containing multiple navigable points and a set of machine offset angles, and outputs a natural language instruction description of the path, employing an encoder-decoder structure based on Transformer or LSTM.
[0014] Furthermore, the input to the trajectory-instruction matcher is a dual-branch structure. One branch consists of a set of environmental visual images containing multiple navigable points and a set of machine offset angles. The other branch is a natural language instruction. The dual-tower structure uses two independent Transformer encoders to encode the two branches separately. The similarity of the output features of the two branches is calculated using cosine similarity, and the calculation formula is as follows:
[0015]
[0016] In the formula, α and β are both normalized vectors of length N, and α = (a1, ..., a2) / (a3, ..., a4) / (a5, ..., a6) / (a7, ..., a8) / (a9, ..., a1) / (a1, ..., a1) / (a2 ... N ),β=(b1,…,b N ).
[0017] Furthermore, in step 3), the candidate navigation points collected in the environment are several discrete points that the machine can actually reach. Several candidate navigation points are connected to form the machine's travel trajectory, with the distance between connected candidate navigation points not exceeding the pre-configured distance as a constraint. Natural language instructions are labeled on the travel trajectory based on the manual annotation method.
[0018] Further, in step 4), the trajectory-instruction generator is trained based on the cross-entropy loss function, which is:
[0019]
[0020] In the formula, N represents the length of the generated text, and f θ (·) represents the output probability of the trajectory-instruction generator model with parameter θ. This represents the i-th truth text.
[0021] Further, in step 4), the trajectory-instruction matcher is trained based on the InfoNCE loss function, whereby the InfoNCE loss function is:
[0022]
[0023] In the formula, M represents the total number of data pairs, < I j ,T j > represents the cosine similarity between text vector I and image vector T, τ represents the temperature vector, and B represents the number of positive and negative samples in each calculation.
[0024] Further, in step 5), the average cosine similarity of the positive examples output on the training set is set as a threshold. The cosine similarity of the output of the dual-tower branches of the trained trajectory-instruction matcher is used to characterize the matching degree between the pseudo-label and the acquisition route. If it is lower than the threshold, the pseudo-label pair is filtered out. The formula for calculating the threshold is as follows:
[0025]
[0026] In the formula, M represents the total number of data pairs, < I i ,T i > represents the cosine similarity between text vector I and image vector T, and τ represents the temperature vector.
[0027] Furthermore, in step 6), when refining the trajectory-instruction generator, the previously trained trajectory-instruction generator weights are loaded, and optimization is performed based on the cross-entropy loss function and by reducing the learning rate value.
[0028] A natural language trajectory instruction generation device based on semi-supervised learning includes a memory, a processor, and a program stored in the memory, wherein the processor executes the program to implement the method described above.
[0029] A storage medium having a program stored thereon, which, when executed, implements the method described above.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] (1) This invention proposes a trajectory-instruction matcher based on a dual-tower structure. It uses adaptive learning to score the matching degree of pseudo-data pairs. If the matching value is less than the threshold, it is filtered out, which ensures the quality of pseudo-labels, effectively reduces the risk of data noise and potential interference with model learning, and improves the accuracy of natural language instruction generation.
[0032] (2) This invention proposes a trajectory-instruction generator based on an encoder-decoder structure. By combining small-scale labeled datasets with large-scale unlabeled datasets using semi-supervised learning, it effectively reduces the reliance on manual annotation and improves the model's accuracy and generalization. The high-performance trajectory-instruction generator obtained by this invention can be used to further generate large-scale, high-quality pseudo-label datasets, providing data support for other similar visual language tasks. Attached Figure Description
[0033] Figure 1 This is a flowchart of the method of the present invention;
[0034] Figure 2 This is a schematic diagram of the trajectory-instruction generator model based on the encoder-decoder structure of the present invention;
[0035] Figure 3 This is a schematic diagram of the trajectory-instruction matcher model based on a dual-tower structure according to the present invention. Detailed Implementation
[0036] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0037] This embodiment provides a method for generating natural language trajectory instructions based on semi-supervised learning, such as... Figure 1 As shown, it includes the following steps:
[0038] Step 1) Construct a trajectory-instruction generator based on an encoder-decoder structure.
[0039] The trajectory-instruction generator takes as input a set of environmental visual images containing multiple navigable points and a set of machine offset angles, and outputs a natural language instruction description of the path, using an encoder-decoder structure based on Transformer or LSTM.
[0040] The Transformer requires significant pre-training, has a large number of parameters, and achieves higher accuracy, making it suitable for applications where high accuracy is required but real-time performance is not critical. LSTM requires no pre-training, has a smaller number of parameters, and lower accuracy, making it suitable for applications where high accuracy is less critical but real-time performance is paramount. When using the Transformer framework, static trigonometric function position encoding or dynamic learnable position encoding can be used to add relative position information to the input. The input to the trajectory-instruction generator includes path image features and machine offset angles, where the machine offset angle is calculated using the following formula:
[0041]
[0042] In the formula, Δ represents the machine offset angle, and θ represents the offset heading angle. This represents the pitch angle of the offset. To avoid a large discrepancy between the dimension of the offset angle and the dimension of the image features, the offset angle can be multiplied by N.
[0043] like Figure 2As shown, the encoder performs behavior encoding and environment encoding based on the features of the input path image, and the decoder uses machine offset angle to add position encoding information to achieve decoding, thereby realizing natural language trajectory instruction prediction.
[0044] Step 2) Construct a trajectory-instruction matcher based on a dual-tower structure.
[0045] The trajectory-instruction matcher has a dual-branch structure as its input. One branch consists of a set of environmental visual images containing multiple navigable points and a set of machine offset angles. The other branch is the natural language instruction. The dual-tower structure uses two independent Transformer encoders to encode the two branches separately. The similarity between the output features of the two branches is calculated using cosine similarity, and the formula is as follows:
[0046]
[0047] In the formula, α and β are both normalized vectors of length N, and α = (a1, ..., a2) / (a3, ..., a4) / (a5, ..., a6) / (a7, ..., a8) / (a9, ..., a1) / (a1, ..., a1) / (a2 ... N ),β=(b1,…,b N ).
[0048] In this embodiment, as Figure 3 As shown, the dual-tower encoders are both based on the Transformer encoder framework, and each projects onto a subspace of the same dimension using a linear mapper. To improve the performance of the matcher, a pre-trained CLIP model is used to extract image features and text features separately. To facilitate the construction of positive and negative samples, trajectory-instruction pairs from the same small batch of training samples are taken as positive examples. Figure 3 Data pairs on the diagonal of the middle slope, together with other data from the same small sample batch, form negative examples. Figure 3 (Data pairs on the diagonal of the China-Africa border). To unify the dimensionality, a soft attention-based approach is used to weight and fuse language instruction embedding features, calculated as follows:
[0049] M = tanh(H)
[0050] α = softmax(MW)
[0051] h = tanh(α) T H)
[0052] In the formula, H represents the language instruction embedding feature with a length of L, tanh represents the tanh activation function, W represents the learnable weights, and h represents the weighted fused feature with a length of 1. This soft attention method can also be applied to the computation of trajectory image features. In particular, the trajectory image branch and the instruction language branch cannot be cross-fused before calculating the cosine similarity, otherwise it will interfere with the actual output of the matcher.
[0053] Step 3) Collect several candidate navigation points in the environment, generate a limited number of trajectories and label them with corresponding natural language instructions. Manually label the data to form a labeled dataset. At the same time, randomly generate a large number of trajectory routes to form an unlabeled dataset.
[0054] Specifically, candidate navigation points collected in the environment are several discrete points that are actually reachable by the machine and are relatively iconic. Following the constraint that the distance between adjacent candidate navigation points does not exceed 3 meters, several candidate navigation points are randomly connected to form the machine's travel trajectory. Natural language instructions are then annotated onto the travel trajectory using a manual annotation method. To improve the generalization of the dataset, each trajectory can correspond to multiple natural language instruction annotations when annotating these trajectories with natural language instructions.
[0055] Step 4) Train the trajectory-instruction generator and the trajectory-instruction matcher using the labeled dataset.
[0056] Both the trajectory-instruction generator and the trajectory-instruction matcher use the gradient descent method to iteratively optimize the internal parameters, and use optimizers such as Adam or SGD to calculate the parameter update method for each time.
[0057] Among them, the trajectory-instruction generator trained based on the cross-entropy loss function is calculated using the following formula:
[0058]
[0059] In the formula, N represents the length of the generated text, and f θ (·) represents the output probability of the trajectory-instruction generator model with parameter θ. This represents the i-th truth text.
[0060] The trajectory-instruction matcher trained based on the InfoNCE loss function is calculated using the following formula:
[0061]
[0062] In the formula, M represents the total number of data pairs. j ,T j > represents the cosine similarity between text vector I and image vector T, τ represents the temperature vector, and B represents the number of positive and negative samples in each calculation.
[0063] Step 5) Based on the unlabeled dataset, generate corresponding pseudo-labels using the trajectory-instruction generator, and filter out low-quality pseudo-labels using the trajectory-instruction matcher.
[0064] The trajectory-instruction generator uses the softmax function to predict the probability of each word:
[0065]
[0066] In the formula, M represents the dictionary dimension size of the natural language instructions, p i d represents the probability that the position is the i-th word. i This represents the numerical value of the i-th word predicted by the trajectory-instruction generator. Using a greedy algorithm, the output at the current position is selected based on the probability of the word with the highest probability.
[0067] To improve the quality of pseudo-labels, the average cosine similarity of the positive examples output on the training set is used as the filtering threshold. The cosine similarity of the output of the dual-tower branches of the trained trajectory-instruction matcher is used to characterize the matching degree between the pseudo-label and the data collection route. If it is lower than the threshold, the pseudo-label pair is filtered out. The formula for calculating the threshold is as follows:
[0068]
[0069] Step 6) Merge the filtered high-quality pseudo-label dataset with the labeled dataset to refine the trajectory-instruction generator and obtain a high-performance trajectory-instruction generator.
[0070] When refining the trajectory-instruction generator, the weights of the previously trained trajectory-instruction generator need to be loaded, and the model is further optimized by using the cross-entropy loss function and reducing the learning rate.
[0071] Step 7) Repeat steps 5) and 6) until the trajectory-instruction matcher determines that there are no low-quality pseudo-labels, or the maximum number of repetition rounds is reached.
[0072] In this process of refinement, the performance of the trajectory-instruction generator gradually improves, and the initially low-quality pseudo-labels can be gradually transformed into high-quality pseudo-labels that pass the trajectory-instruction matcher's evaluation. This enhances the effectiveness of supervised learning and continuously improves the model's generalization and accuracy. Typically, this process is repeated 2-3 times.
[0073] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0074] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for generating natural language trajectory instructions based on semi-supervised learning, characterized in that, Includes the following steps: Step 1) Construct a trajectory-instruction generator based on an encoder-decoder structure; Step 2) Construct a trajectory-instruction matcher based on a dual-tower structure; Step 3) Collect several candidate navigation points in the environment, generate a limited number of trajectories and label them with corresponding natural language instructions. Manually label the data to form a labeled dataset. At the same time, randomly generate a large number of trajectory routes to form an unlabeled dataset. Step 4) Train the trajectory-instruction generator and the trajectory-instruction matcher using the labeled dataset; Step 5) Based on the unlabeled dataset, generate corresponding pseudo-labels using the trajectory-instruction generator, and filter out low-quality pseudo-labels using the trajectory-instruction matcher. Step 6) Merge the filtered high-quality pseudo-label dataset with the labeled dataset to refine the trajectory-instruction generator and obtain a high-performance trajectory-instruction generator; Step 7) Repeat steps 5) and 6) until the trajectory-instruction matcher determines that there are no low-quality pseudo-labels, or the maximum number of repetition rounds is reached; The trajectory-instruction matcher has a dual-branch structure as its input. One branch consists of a set of environmental visual images containing multiple navigable points and a set of machine offset angles. The other branch is a natural language instruction. The dual-tower structure uses two independent Transformer encoders to encode the two branches separately. The similarity between the output features of the two branches is calculated using cosine similarity, and the calculation formula is as follows: In the formula, and All are normalized lengths of The vector, ; In step 5), the average cosine similarity of the positive examples output on the training set is set as the threshold. The cosine similarity of the output of the dual-tower branches of the trained trajectory-instruction matcher is used to characterize the matching degree between the pseudo-label and the acquisition route. If it is lower than the threshold, the pseudo-label pair is filtered out. The formula for calculating the threshold is as follows: In the formula, Indicates the total number of data pairs. Representing text vectors and image vectors Cosine similarity between them This represents the temperature coefficient.
2. The method for generating natural language trajectory instructions based on semi-supervised learning according to claim 1, characterized in that, The trajectory-instruction generator takes as input a set of environmental visual images containing multiple navigable points and a set of machine offset angles, and outputs a natural language instruction description of the path, employing an encoder-decoder structure based on Transformer or LSTM.
3. The method for generating natural language trajectory instructions based on semi-supervised learning according to claim 1, characterized in that, In step 3), the candidate navigation points collected in the environment are several discrete points that the machine can actually reach. Several candidate navigation points are connected to form the machine's travel trajectory, with the distance between connected candidate navigation points not exceeding the pre-configured distance as a limit. Natural language instructions are labeled on the travel trajectory based on the manual annotation method.
4. The method for generating natural language trajectory instructions based on semi-supervised learning according to claim 1, characterized in that, In step 4), the trajectory-instruction generator is trained based on the cross-entropy loss function, which is: In the formula, Indicates the length of the generated text. The parameter is The output probability of the trajectory-instruction generator model, Indicates the first i A truth text.
5. The method for generating natural language trajectory instructions based on semi-supervised learning according to claim 1, characterized in that, In step 4), the trajectory-instruction matcher is trained based on the InfoNCE loss function, which is: In the formula, Indicates the total number of data pairs. Representing text vectors and image vectors Cosine similarity between them This indicates the number of positive and negative samples in each calculation.
6. The method for generating natural language trajectory instructions based on semi-supervised learning according to claim 1, characterized in that, In step 6), when refining the trajectory-instruction generator, the weights of the previously trained trajectory-instruction generator are loaded, and the learning rate value is reduced based on the cross-entropy loss function for optimization.
7. A natural language trajectory instruction generation device based on semi-supervised learning, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-6.
8. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-target tracking unsupervised domain adaptation method based on pseudo label correction
CN114693979A
Semi-supervised method for strip mine card state identification under time sequence GAN data enhancement
CN115130599A