A visual language navigation method combining image description and text generation
By combining image description and text generation methods to generate images similar to the current scene, the problem of data scarcity in visual language navigation is solved, and task performance and model generalization capabilities are improved.
Patent Information
- Application Number
- CN202311520271.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-11-15
AI Technical Summary
Existing visual language navigation technology faces the problem of data scarcity, especially in obtaining enough real three-dimensional environmental data with a height viewpoint of the human body, which leads to the scarcity of available scenes in the simulator, which hinders the development of visual language navigation.
Using a method of combining image description and text generation images, detailed natural language image descriptions are generated through the scene description module, and a text generation image model is used to generate scene images similar to the current scene as additional visual data input for visual language navigation.
By generating images similar to the current scene, additional visual data input is provided, which improves the task performance of visual language navigation agents and generalizes the model, and overcomes the problem of data scarcity.
Smart Images

Figure CN117571014B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of visual language navigation, and relates to a visual language navigation method for generating an image by combining image description and text. Background Art
[0002] In the field of artificial intelligence, natural language processing and computer vision technologies have achieved remarkable results in multiple tasks. Due to their rapid development, an increasingly strong trend is to combine them to achieve multimodal machine learning to handle more comprehensive tasks such as visual question answering, image description, and text-to-image generation. Among them, visual language navigation, as an important branch of multimodal machine learning and embodied intelligence, has attracted widespread attention. On the basis of having both visual and language capabilities, this task further requires the agent to be able to make action decisions based on current information.
[0003] The visual language navigation task requires the agent to combine natural language instructions and first-person visual information to be able to autonomously navigate in an unfamiliar environment to achieve the goals set by the instructions. For example, in an indoor home environment, people may need the help of an agent and send simple instructions such as "help me go to the bedroom and find the book on the bedside table." This requires the visual language navigation agent to have the ability to understand natural language instructions and observe the surrounding environment, and to make inference decisions based on the visual and language information obtained, and finally complete the navigation task.
[0004] It can be seen that the visual language navigation task is a very challenging and comprehensive task that requires consideration of multiple information sources, including panoramic visual information, language instruction information, historical decision information, and historical visual information. However, for such a complex task, current research mostly relies on simulated environments built based on real scene images, and the construction of such environments is usually time-consuming and laborious. Among them, it is most difficult to obtain enough real three-dimensional stereoscopic environment data with human height viewpoints. For example, the Matterprt3D simulator only contains 60 different room environments for intelligent agents to train, and the remaining environments need to be used for verification and testing. Although many studies have made efforts to enrich the amount of data from the perspective of increasing the number of instructions, it is still difficult to avoid the problem of the scarcity of available scenes in the simulator, so the problem of data scarcity has always been one of the main challenges hindering the development of visual language navigation. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a visual language navigation method that combines image description and text to generate images. To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A visual language navigation method for generating images by combining image description and text, the method comprising the following steps:
[0007] S1: Obtain the natural language target instructions of the visual language navigation task and the visual image of the current scene location;
[0008] S2: Based on the panoramic visual image obtained in S1, a scene description module is used to generate a detailed natural language image description of the room type, core objects, the relationship between core objects, and the core scene layout of the current scene;
[0009] S3: Use the detailed natural language image description generated in S2 as the input of the text generation image model, and finally generate a similar scene with similar core objects and core scene layout as the described scene;
[0010] S4: The visual image in S1 and the similar scene image generated based on the current scene in S3 are respectively extracted with visual features through a multi-layer Transformer structure. At the same time, the natural language target instruction in S1 is passed through a text encoder to obtain the target instruction encoding. Then, the two extracted visual features are encoded through a fine-scale cross-modal encoder in combination with the target instruction encoding, and finally the current scene encoding and the similar scene encoding are generated.
[0011] S5: The current scene encoding and similar scene encoding generated in S4 are used to generate visually enhanced scene fusion features through the cross attention layer, and then injected into the linear feedforward network. Then, the visually enhanced action prediction for the next execution action is generated based on all current waypoints through the Softmax activation function. The visually enhanced action prediction and the benchmark expert action are cross-entropy operated to generate enhanced action loss, that is, the learning of visually enhanced scene fusion features is guided by supervising visually enhanced action prediction, which is described by the formula:
[0012]
[0013]
[0014]
[0015] Where t represents the current time step, represents the visual enhancement scene fusion feature, Cross-Attn represents the cross attention layer, Indicates the current scene encoding. Indicates similar scene encoding, represents visually enhanced action prediction, FFN represents linear feedforward network, represents the enhanced action prediction loss, CrossEntropy represents the cross entropy loss function, represents the benchmark expert action;
[0016] S6: Furthermore, at each time step, the current scene encoding generated in S4 and the visually enhanced scene fusion features generated in S5 are aggregated through a linear feedforward network and a Sigmoid activation function to generate dynamic fusion weights for dynamically fusion of the visually enhanced action prediction in S5 and the action prediction based on the current scene:
[0017]
[0018] Among them, σ t represents the learnable dynamic fusion weight, based on which the final navigation decision is expressed as:
[0019]
[0020] in, It means that the fused action prediction of the current real scene and the corresponding similar scene is comprehensively considered, and finally the cross entropy calculation is performed between the fused action prediction and the benchmark expert action:
[0021]
[0022] in, represents the fused action prediction loss, which guides the learning of the entire decision process by supervising the fused action prediction.
[0023] Optionally, in S2, the specific process of generating a detailed natural language image description through the scene description module is:
[0024] First, the panoramic image of the current scene is discretized into 36 first-person visual images by uniformly adjusting the agent's perspective. Then, for each visual image, a pre-trained visual-language model CLIP or BILP-2 is used to give an overview of the current image and generate an overview of the current room type and the core objects in the image. At the same time, based on the open source dataset Visual Genome A scene description corpus is established for specific scenes, such as indoor home scenes, in the format of "attribute-object" two-tuples and "subject-predicate-object" triples. The current image is then cut into five sub-images at the upper left, upper right, lower left, lower right and common center image positions. Based on the multimodal pre-trained model CLIP, its text encoder CLIP-T is used to encode the entire scene description corpus as a search keyword. Then, its image encoder CLIP-I is used to encode each of the cut sub-images respectively and use it as a query to find the keyword with the highest cosine similarity, that is, the highest matching degree, in the scene description corpus. The scene description corresponding to the keyword is used as the description of the current sub-image, forming a detailed description of the core object or the relationship between the core object objects in the sub-image. In summary, a general overview and five detailed descriptions can be obtained in the end. They are combined in the order of general overview first and then detailed description, and combined with the corresponding positional relationship of the detailed description to generate a detailed natural language image description.
[0025] Optionally, in S3, the text-to-image model refers to an advanced model for the text-to-image task in the field of multimodal machine learning, including Stable Diffusion.
[0026] Optionally, in S4, the text encoder encodes each word to represent the position code and word type code of the word relative to the entire sentence, and finally injects the position code and word type code together into the multi-layer Transformer structure.
[0027] Optionally, in S4, the fine-scale cross-modal encoder is specifically an image feature extraction network, including ResNet, ViT or CLIP, a cross-attention layer that integrates target instruction language features and visual image features, a self-attention layer, and a linear feedforward network.
[0028] The beneficial effect of the present invention is that, inspired by the fact that humans rely on experience to actively associate similar scenes they have seen before to help make behavioral decisions in unfamiliar scenes, by combining advanced models of image description and text generation image tasks in the field of multimodal machine learning, similar scenes are generated based on the current scene, providing additional visual data input for task training, thereby improving the task performance of the visual language navigation agent and the generalization ability of the model.
[0029] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0031] Figure 1 The figure is a flow chart of the method of the present invention.
[0032] Figure 2 It is the overall framework structure diagram of the method of the present invention. DETAILED DESCRIPTION
[0033] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0034] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0035] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0036] See also Figure 1 and Figure 2 , the present invention discloses a visual language navigation method for generating images by combining image description and text, and the detailed steps are as follows:
[0037] S1: Obtain the natural language target instructions of the visual language navigation task and the panoramic visual image of the current scene. Based on the three directions of looking up, looking straight and looking down and the 12 directions of horizontal rotation of the agent, the current panoramic visual image is discretized into 36 first-person visual images by evenly adjusting the agent's perspective;
[0038] S2: For each first-person visual image generated in S1, a detailed natural language image description is generated through scene description. The specific implementation process of scene description is as follows: first, the visual language pre-trained model BILP-2 is used to give an overall overview of the current image, and an overview of the current room type and the core objects in the image is generated. At the same time, based on the open source dataset Visual Genome Dataset, a scene description corpus is established for specific scenes, such as indoor home scenes, in the format of "attribute-object" two-tuples and "subject-predicate-object" triples. The current image is then cut into five sub-images at the upper left, upper right, lower left, lower right and common center image positions. Based on the multimodal pre-trained model CLIP, its text encoder CLIP-T is used to encode the entire scene description corpus as a search keyword, and then its image encoder CLIP-I is used to encode each of the cut sub-images and use it as a query to find the keyword with the highest cosine similarity, that is, the highest matching degree, in the scene description corpus. The scene description corresponding to the keyword is used as the description of the current sub-image, forming a detailed description of the core object or the relationship between the core object in the sub-image. In summary, a general overview and five detailed descriptions can be obtained in the order of general overview first and then detailed description, and combined with the corresponding positional relationship of the detailed description to generate a detailed natural language image description describing the room type, core object, relationship between core object objects and core scene layout of the current scene location;
[0039] S3: The detailed natural language image description generated in S2 is used as the input of Stable Diffusion, an advanced model for text-to-image generation tasks, to generate similar scenes with similar core objects and core scene layouts as the described scene;
[0040] S4: The visual features of the visual image in S1 and the similar scene image generated based on the current scene in S3 are extracted through a multi-layer Transformer structure respectively. At the same time, the natural language target instruction in S1 is encoded, and additional position encoding and word type encoding representing the word relative to the entire sentence are added. Then they are injected into the multi-layer Transformer structure together to generate the target instruction encoding. The two extracted visual features are encoded through a fine-scale cross-modal encoder in combination with the target instruction encoding. The fine-scale cross-modal encoder includes an image feature extraction network CLIP, a cross-attention layer that integrates the target instruction language features and visual image features, a self-attention layer, and a linear feedforward network composed of a fully connected layer, a Relu activation function, a Bert normalization layer, and a fully connected layer. Finally, the current scene encoding and the similar scene encoding are generated;
[0041] S5: For the current time step t, encode the current scene generated in S4 and similar scene coding Generating visually enhanced scene fusion features via criss-cross attention layers
[0042]
[0043] Among them, Cross-Attn represents the cross attention layer, and then the visual enhancement scene fusion features are combined based on all current navigable points Injected into the linear feedforward network, and then generated through the Softmax activation function to generate visually enhanced action predictions for the next action to be performed
[0044]
[0045] Among them, FFN represents a linear feedforward network, and finally the visual enhanced action prediction and benchmark expert actions Perform cross entropy operation to generate enhanced action loss
[0046]
[0047] Among them, CrossEntropy represents the cross entropy loss function, which enhances action prediction through supervised vision To guide the visual enhancement scene fusion feature learning;
[0048] S6: Furthermore, for each time step t, the current scene encoding generated in S4 is aggregated through a linear feed-forward network FFN and a Sigmoid activation function Fusion features with the visually enhanced scene generated in S5 Vision-enhanced action prediction for dynamic fusion S5 And generate dynamic fusion weights based on the action prediction made in the current scene:
[0049]
[0050] Among them, σ t represents the learnable dynamic fusion weight, based on which the final navigation decision is expressed as:
[0051]
[0052] in, It means that the fusion action prediction of the current real scene and the corresponding similar scene is comprehensively considered, and finally the fusion action prediction is Actions with Benchmark Expert Perform cross entropy calculation:
[0053]
[0054] in, represents the fusion action prediction loss, which is obtained by supervising the fusion action prediction To guide learning throughout the decision-making process.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A visual language navigation method for generating images by combining image description and text, characterized in that: The method comprises the following steps: S1: Obtain the natural language target instructions of the visual language navigation task and the visual image of the current scene location; S2: Based on the panoramic visual image obtained in S1, a scene description module is used to generate a detailed natural language image description of the room type, core objects, the relationship between core objects, and the core scene layout of the current scene; S3: Use the detailed natural language image description generated in S2 as the input of the text generation image model, and finally generate a similar scene with similar core objects and core scene layout as the described scene; S4: The visual image in S1 and the similar scene image generated based on the current scene in S3 are respectively extracted with visual features through a multi-layer Transformer structure. At the same time, the natural language target instruction in S1 is passed through a text encoder to obtain the target instruction encoding. Then, the two extracted visual features are encoded through a fine-scale cross-modal encoder in combination with the target instruction encoding, and finally the current scene encoding and the similar scene encoding are generated. S5: The current scene encoding and similar scene encoding generated in S4 are used to generate visually enhanced scene fusion features through the cross attention layer, and then injected into the linear feedforward network. Then, the visually enhanced action prediction for the next execution action is generated based on all current waypoints through the Softmax activation function. The visually enhanced action prediction and the benchmark expert action are cross-entropy operated to generate enhanced action loss, that is, the learning of visually enhanced scene fusion features is guided by supervising visually enhanced action prediction, which is described by the formula: Where t represents the current time step, represents the visual enhancement scene fusion feature, Cross-Attn represents the cross attention layer, Indicates the current scene encoding. Indicates similar scene encoding, represents visually enhanced action prediction, FFN represents linear feedforward network, represents the enhanced action prediction loss, CrossEntropy represents the cross entropy loss function, represents the benchmark expert action; S6: Furthermore, at each time step, the current scene encoding generated in S4 and the visually enhanced scene fusion features generated in S5 are aggregated through a linear feedforward network and a Sigmoid activation function to generate dynamic fusion weights for dynamically fusion of the visually enhanced action prediction in S5 and the action prediction based on the current scene: Among them, σ t represents the learnable dynamic fusion weight, based on which the final navigation decision is expressed as: in, It means that the fused action prediction of the current real scene and the corresponding similar scene is comprehensively considered, and finally the cross entropy calculation is performed between the fused action prediction and the benchmark expert action: in, represents the fused action prediction loss, which guides the learning of the entire decision process by supervising the fused action prediction.
2. The visual language navigation method for generating images by combining image description and text according to claim 1, characterized in that: In S2, the specific process of generating a detailed natural language image description through the scene description module is as follows: First, the panoramic image of the current scene is discretized into 36 first-person visual images by evenly adjusting the agent's perspective. Then, for each visual image, a pre-trained multimodal pre-trained model CLIP or BILP-2 is used to give an overview of the current image, generating an overview of the current room type and the core objects in the image. At the same time, based on the open source dataset Visual Genome The scene description corpus is established for specific scenes, including indoor home scenes, in the format of "attribute-object" two-tuples and "subject-predicate-object" triples. The current image is then cut into five sub-images at the upper left, upper right, lower left, lower right and common center image positions. Based on the multimodal pre-trained model CLIP, its text encoder CLIP-T is used to encode the entire scene description corpus as a search keyword. Then, its image encoder CLIP-I is used to encode each of the cut sub-images and use it as a query to find the keyword with the highest cosine similarity, that is, the highest matching degree, in the scene description corpus. The scene description corresponding to the keyword is used as the description of the current sub-image, forming a detailed description of the core object or the relationship between the core object objects in the sub-image. In summary, a general overview and five detailed descriptions can be obtained in the end. They are combined in the order of general overview first and detailed description, and combined with the corresponding positional relationship of the detailed description to generate a detailed natural language image description.
3. The visual language navigation method for generating images by combining image description and text according to claim 1, characterized in that: In S3, the text-to-image model refers to an advanced model for the text-to-image task in the field of multimodal machine learning, including Stable Diffusion.
4. The visual language navigation method for generating images by combining image description and text according to claim 1, characterized in that: In S4, the text encoder encodes each word to represent the position code and word type code of the word relative to the entire sentence, and finally injects the position code and word type code together into the multi-layer Transformer structure.
5. The visual language navigation method for generating images by combining image description and text according to claim 1, characterized in that: In S4, the fine-scale cross-modal encoder includes an image feature extraction network, a cross-attention layer, a self-attention layer and a linear feedforward network that fuses target instruction language features and visual image features; wherein the image feature extraction network includes ResNet, ViT and CLIP.
Citation Information
Patent Citations
Image generation method and terminal device
CN110136216A
Target-driven visual semantic navigation method in indoor scene, storage medium and equipment
CN115439728A