Generative biological incentive design method driven by visual language model

Through the generative bio-stimulus design method driven by visual language model, combined with search engines and diffusion models, the problem of insufficient efficiency and intelligence in the existing technology is solved, and efficient and intelligent design generation and optimization are achieved, with a wide range of interdisciplinary application potential.

CN120069062APending Publication Date: 2025-05-30FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105506.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology has shortcomings in bionic design efficiency, degree of intelligence, cross-field applications, etc., and it is difficult to efficiently and intelligently generate innovative design solutions that meet biological characteristics.

Method used

The generative biological excitation design method driven by visual language model is adopted, combining search engines' biological domain instance retrieval, prompt word generation and diffusion model image generation capabilities of visual language models to achieve efficient and intelligent design generation and optimization.

Benefits of technology

It significantly improves design efficiency and solution quality, reduces manual intervention, and achieves cross-disciplinary and cross-field design innovation. The generated design solutions are highly innovative, functionally adaptable and practical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069062A_ABST
    Figure CN120069062A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language model-driven generative biological incentive design method. The method comprises the following steps of: inputting an engineering field design problem by a user; obtaining a biological domain instance corresponding to the search word based on a search engine; the user cooperates with the visual language model to generate cue words; inputting cue words to the diffusion model to generate a plurality of preliminary design schemes; the user preliminarily evaluates the design scheme and judges whether the design scheme needs to be regenerated or not; if the user is not satisfied with the generated design scheme, the user cooperates with the visual language model to adjust and optimize the cue word, and the design scheme is regenerated; and performing quantitative evaluation on the design scheme based on the visual language model, and selecting a final design scheme. According to the method, the visual language model and the diffusion model are combined, the innovative design scheme can be rapidly generated, the design scheme is intelligently evaluated and optimized through the fine-adjusted visual language model, the design efficiency and the scheme quality are remarkably improved, and manual intervention is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent design and bionic technology, and in particular to a visual language model-driven generative bio-inspired design method. Background Art

[0002] Bionic design is a design method that draws on the structure, function, behavior and evolutionary characteristics of organisms in nature to create innovative products that meet engineering needs. In recent years, bionic design has gradually become an important direction in the field of engineering design and is widely used in mechanical design, robotics, architectural design, and material science. However, traditional bionic design methods mostly rely on the experience and intuition of designers, which makes the design process time-consuming and has certain limitations. Traditional methods are difficult to efficiently extract valuable information from a large number of biological examples, and they are also unable to flexibly combine biological characteristics with design requirements for systematic design.

[0003] In recent years, with the rapid development of artificial intelligence technology, especially the widespread application of large language models (such as ChatGPT) and generative models (such as diffusion models), the potential of artificial intelligence in bionic design has become increasingly apparent. Large language models can process complex text information and generate high-quality language output by deeply learning large-scale corpora; while diffusion models can perform well in image generation and have strong creativity and controllability. The combination of these technologies provides a new methodology for bionic design, which can efficiently and systematically generate design solutions that conform to biological characteristics, and intelligently evaluate and optimize them.

[0004] Although artificial intelligence technology has broad application prospects in the field of bionic design, how to effectively combine large language models with diffusion models and use massive data and multimodal information of biological examples for intelligent design is still a challenge that needs to be solved. Most existing technologies focus on a single modality or a specific field, lack interdisciplinary and cross-domain versatility, and often face problems such as low efficiency and poor innovation in actual design. Therefore, how to build an efficient, intelligent, and innovative generative bionic design method has become a technical problem that needs to be overcome in the field of bionic design.

[0005] The present invention is proposed to address the shortcomings of the prior art in terms of bionic design efficiency, intelligence, cross-domain application, etc., and to provide a generative bio-inspired design method driven by a visual language model. Summary of the invention

[0006] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a generative bio-inspired design method driven by a vision-language model. This method combines the ability of a search engine to retrieve examples in the biological field, the ability of a vision-language model to generate prompt words, and the ability of a diffusion model to generate images, and can efficiently and intelligently generate innovative design solutions that conform to biological characteristics, and comprehensively evaluate and optimize these design solutions.

[0007] The present invention provides the following technical solutions:

[0008] The present invention provides a generative bio-inspired design method driven by a vision-language model, including the following steps:

[0009] S01. The user inputs a design problem in the engineering field;

[0010] S02. Based on a search engine, obtain biological field examples corresponding to the search terms;

[0011] S03. The user collaborates with the vision-language model to generate prompt words;

[0012] S04. Input the prompt words into the diffusion model to generate multiple preliminary design solutions;

[0013] S05. The user makes a preliminary evaluation of the design solutions and determines whether to regenerate;

[0014] S06. If the user is not satisfied with the generated design solutions, the user collaborates with the vision-language model to adjust and optimize the prompt words and regenerate the design solutions;

[0015] S07. Based on the vision-language model, conduct a quantitative evaluation of the design solutions and select the final design solution.

[0016] As a preferred technical solution of the present invention, the user inputting a design problem in the engineering field in step S01 includes a text description of the design problem.

[0017] As a preferred technical solution of the present invention, the search engine in step S02 includes BARCODE.

[0018] As a preferred technical solution of the present invention, the vision-language model in step S03 includes GPT-4V; the user collaborating with the vision-language model to generate prompt words in step S03 includes the user generating prompt words, the vision-language model generating prompt words, and the user collaborating with the vision-language model to generate prompt words; the user collaborating with the vision-language model to generate prompt words in step S03 includes positive prompt words and negative prompt words.

[0019] As a preferred technical solution of the present invention, the diffusion model in step S04 includes StableDiffusion 3; generating multiple preliminary design schemes in step S04 includes picture descriptions of the design schemes; generating multiple preliminary design schemes in step S04 includes dozens, hundreds, thousands or even more design schemes.

[0020] As a preferred technical solution of the present invention, the user's preliminary evaluation of the design scheme in step S05 includes subjective evaluation of the design scheme.

[0021] As a preferred technical solution of the present invention, the user and the vision-language model cooperate to adjust and optimize the prompt words in step S06, including the user's adjustment and optimization of the prompt words, the vision-language model's adjustment and optimization of the prompt words, and the user and the vision-language model's cooperation to adjust and optimize the prompt words; the user and the vision-language model cooperate to adjust and optimize the prompt words in step S06, including the adjustment and optimization of positive prompt words, the adjustment and optimization of negative prompt words, and the adjustment and optimization of both positive and negative prompt words simultaneously.

[0022] As a preferred technical solution of the present invention, the quantitative evaluation of the design scheme based on the vision-language model in step S07 includes fine-tuning the vision-language model; the quantitative evaluation of the design scheme based on the vision-language model in step S07 includes that the evaluation indicators are innovativeness, aesthetics, structural rationality, and functional adaptability.

[0023] As a preferred technical solution of the present invention, this method is widely applicable to innovative designs in fields such as bionic robots, industrial equipment, and intelligent manufacturing.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] 1: Efficient and intelligent design generation and optimization: By combining the vision-language model and the diffusion model, the present invention can quickly generate a large number of innovative design schemes, and use the fine-tuned vision-language model to conduct intelligent evaluation and optimization of the design schemes, significantly improving the design efficiency and the quality of the schemes, and reducing manual intervention.

[0026] 2. Precise guidance and interdisciplinary applications: The positive and negative prompt words generated by the vision-language model can precisely guide the extraction of biological features and target matching in the design, ensuring that the design results conform to the biological characteristics of nature. This method has broad interdisciplinary application potential and can be widely applied to multiple fields such as bionic robots, intelligent equipment, and materials science.

[0027] 3. Innovation and Feasibility Enhancement: The design solutions generated by the present invention have significant advantages in terms of innovation, functional adaptability, and practicality. The multi-resolution image generation ability of the diffusion model ensures that the designs are more creative and highly feasible, promoting the intelligent and systematic development of bionic design methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0029] Figure 1 is a flowchart of a generative bio-inspired design method driven by a vision-language model proposed by the present invention;

[0030] Figure 2 is an example of the retrieval result of biological field instances obtained by the search engine;

[0031] Figure 3 is an example of the prompt words generated by the vision-language model;

[0032] Figure 4 is an example of the design solution generated by the diffusion model;

[0033] Figure 5 is an example of the picture for fine-tuning;

[0034] Figure 6 is an example of the evaluation result of the design solution using the vision-language model; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.

[0036] As Figures 1 - 6 shown:

[0037] Embodiment 1:

[0038] A generative bio-inspired design method driven by a vision-language model includes the following steps:

[0039] S01. The user inputs a design problem in the engineering field. The description of the design problem in step S01 includes text descriptions, such as robotic gripper, etc.;

[0040] S02. Based on the search engine, obtain biological field instances corresponding to the search terms. The search engine in step S02 includes BARCODE. The obtained biological field instances include the names and feature descriptions of the biological field instances, and the description methods include text descriptions;

[0041] S03. The user collaborates with a vision - language model to generate prompts. The vision - language model in step S03 includes GPT - 4V. The ways of generating prompts include user - generated prompts, vision - language - model - generated prompts, and user - vision - language - model collaborative generation of prompts. The generated prompts include positive prompts and negative prompts. The description ways of prompts include text descriptions.

[0042] S04. Input the prompts into a diffusion model to generate multiple preliminary design schemes. The diffusion model in step S04 includes Stable Diffusion 3. The generated design schemes include picture descriptions of the design schemes. The number of generated design schemes can reach dozens, hundreds, thousands or more, and corresponding numbers of design schemes can be generated according to the user's requirements.

[0043] S05. The user makes a preliminary evaluation of the design scheme to judge whether to regenerate. The user's preliminary evaluation of the design scheme in step S05 includes subjective evaluation of the design scheme.

[0044] S06. If the user is not satisfied with the generated design scheme, the user collaborates with the vision - language model to adjust and optimize the prompts and regenerate the design scheme. The user - vision - language - model collaborative adjustment and optimization of the prompts in step S06 include user - adjusted and optimized prompts, vision - language - model - adjusted and optimized prompts, and user - vision - language - model collaborative adjustment and optimization of the prompts. The adjustment and optimization of the prompts in step S06 include adjustment and optimization of positive prompts, adjustment and optimization of negative prompts, and simultaneous adjustment and optimization of positive and negative prompts.

[0045] S07. Based on the vision - language model, quantitatively evaluate the design scheme and select the final design scheme. The vision - language model in step S07 is a fine - tuned vision - language model. The fine - tuning pictures come from the quantitative scoring of pictures by experts. The evaluation indicators of expert scoring include innovation, aesthetics, structural rationality, and functional adaptability. The user can screen the design schemes according to the evaluation results and select the optimal design scheme from them.

[0046] Example 2:

[0047] Based on Example 1, taking the robotic gripper as an example, in BARCODE, the input engineering vocabulary is “robotic gripper”, that is, let query_list = ['robotic gripper']. Initially, consider mining the top 50 biological prototypes first, that is, let top_n_results = 50. Figure 2Shows 5 retrieved biological domain examples corresponding to robotic gripper, sorted according to relevance. For example, the first-ranked biological domain example is Caribbean hermit crab, and the description of this example is "Cittarium pica shell is often used for its home, and the hermit crab can use its larger claw to cover the aperture of the shell for protection against predators."

[0048] For the biological domain examples obtained by BARCODE, Figure 3Taking the Ghost bat as an example, it shows the positive and negative prompts generated by the user in collaboration with the visual language model. For example, the positive prompt for the Ghost bat is "Design a robotic gripper inspired by bat’s claw. The details of the image need to be smooth and delicate. Mimicking the ability of its claw to grab prey. The gripper should combine biomechanical and robotic elements, showcasing metal joints and actuators to replicate the flexibility and precision of a claw. The overall appearance should be realistic, with smooth metal finishes and detailed carvings to emphasize the beauty of the bat's claw. The focus should be entirely on the gripper. The overall design pursues a minimalist style." The negative prompt is "Don't let the gripper become twisted. Don't look like a bat. No animals. No anime style. The subject doesn’t need to be too complicated. No human arms."

[0049] Based on the prompts generated in the above steps, Figure 4 it shows some conceptual design diagrams of the robotic gripper generated by the diffusion model. There are a total of 60 conceptual design diagrams here. The diffusion model can generate dozens, hundreds, thousands or even more design schemes according to the user's instructions.

[0050] Regarding the design schemes generated by the diffusion model, Figure 5Select some of the design solutions for fine-tuning the vision language model. Here, there are a total of 80 design solutions for fine-tuning and their scores. The scores of these design solutions are obtained by experts' scoring. Generally, the number of design solutions for fine-tuning is preferably 200, and it can also be extended to a larger number according to needs. After completing the manual scoring by experts, these image-score pairs are compiled into a dataset file in a fixed format. Each item in the dataset is divided into three parts: system instruction, user question, and VLM's answer. Then, the GPT-4V model is fine-tuned through the API of OpenAI.

[0051] Based on the fine-tuned vision language model, Figure 6 The evaluation scores of some design solutions are shown. For example, the evaluation scores of the 5 design solutions in the first row are 96.15, 95.50, 91.67, 90.30, and 87.45 respectively. Users can select better design solutions based on the evaluation results of the design solutions.

[0052] The present invention provides a brand-new bio-inspired design process. Through intelligent data retrieval, design generation, and evaluation and optimization processes, it not only improves the design efficiency but also realizes cross-disciplinary and cross-field design innovation.

[0053] In summary, the present invention generates positive and negative prompt words through the vision language model, precisely guides the matching of biometric feature extraction and design goals in the design process, and thus realizes the effective combination of biometric features and design language. On the other hand, through the diffusion model, design solutions related to design problems are generated, realizing the automatic generation and batch generation of design solutions. This process can break through the limitations of traditional bionic design methods based on manual experience, reduce the interference of subjectivity in the design process, and improve the scientificity and automation of design.

[0054] First of all, the present invention retrieves bio-domain instance information related to the target design task through the BARCODE search engine. BARCODE can automatically extract bio-inspiration from the network and provide multi-dimensional bio-instance support for the design process. This bio-instance information will become the basis for the present invention to generate bio-inspired design solutions.

[0055] Secondly, based on the retrieved biological instances, the vision-language model generates positive prompt words and negative prompt words by analyzing and processing biological features. The positive prompt words describe the requirements of the target biological features for the design solution, such as the structure, function, adaptability, etc. of the organism; while the negative prompt words reflect the biological features to be avoided in the design or the design requirements that are not suitable, such as some features that do not conform to the laws of biological evolution. By combining these two types of prompt words, the present invention can provide precise design guidance for the generative design model. During this process, the user can adjust and modify the prompt words generated by the vision-language model and collaborate with the vision-language model to complete the generation of the prompt words.

[0056] Then, the above-mentioned prompt words are input into the diffusion model (Stable Diffusion 3), and the diffusion model is used to generate a design solution. Based on the multi-resolution image generation technology, the diffusion model can create a design solution that conforms to the target biological features at different levels of detail. The powerful image generation ability of this model enables the design solution to reflect the evolutionary characteristics and adaptability of the organism in terms of structure, function, and form, effectively breaking through the creative bottleneck in traditional design methods.

[0057] Immediately afterwards, the user conducts a preliminary evaluation of the generated design solution. If the user is not satisfied with the generated design solution, the user and the vision-language model collaborate to adjust and optimize the prompt words until a satisfactory design solution is generated.

[0058] Finally, the generated design solution will be intelligently evaluated through the fine-tuned vision-language model. By using a certain number of bionic design samples for supervised learning of the vision-language model, the present invention can comprehensively evaluate the design solution from multiple dimensions such as innovation, aesthetics, structural rationality, and functional adaptability. The fine-tuned model will conduct multi-dimensional evaluation and optimization of the design solution according to the actual design requirements, so as to screen out the optimal design solution. This process greatly improves the intelligent level of the design solution evaluation and ensures the efficiency and high quality of the design solution evaluation.

[0059] The present invention provides a brand-new bio-inspired design method, which not only has broad application prospects in the field of mechanical design, but also provides reliable technical support for innovative designs in fields such as intelligent equipment, construction engineering, and robotics. By combining the vision-language model and the diffusion model, the present invention can break through the limitations of traditional design methods and achieve cross-disciplinary and cross-field intelligent design, with great commercial and scientific research value.

[0060] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A visual language model driven generative bio-inspired design method, characterized in that: The following steps are involved: S01, the user inputs the design problem in the engineering field; S02. Obtain biological field examples corresponding to the search terms based on the search engine; S03, the user collaborates with the visual language model to generate prompt words; S04, inputting prompt words into the diffusion model to generate multiple preliminary design solutions; S05. The user makes a preliminary assessment of the design solution to determine whether it needs to be regenerated; S06. If the user is not satisfied with the generated design solution, the user collaborates with the visual language model to adjust and optimize the prompt words and regenerate a design solution; S07. Quantitatively evaluate the design schemes based on the visual language model and select the final design scheme.

2. The visual language model-driven generative bio-inspired design method according to claim 1, characterized in that: In step S01, the user inputs a design problem in the engineering field, including a text description of the design problem.

3. The visual language model-driven generative bio-inspired design method according to claim 1, characterized in that: The search engine in step S02 includes BARCODE.

4. The visual language model-driven generative bio-inspired design method according to claim 1, characterized in that: The visual language model in step S03 includes GPT-4V; the user and the visual language model collaborate to generate prompt words in step S03, including the user generating prompt words, the visual language model generating prompt words, and the user and the visual language model collaborate to generate prompt words; the user and the visual language model collaborate to generate prompt words in step S03, including forward prompt words and reverse prompt words.

5. The visual language model driven generative bio-inspired design method according to claim 1, characterized in that: The diffusion model in step S04 includes Stable Diffusion 3; the generation of multiple preliminary design schemes in step S04 includes picture descriptions of the design schemes; the generation of multiple preliminary design schemes in step S04 includes dozens, hundreds, or even thousands or more design schemes.

6. The visual language model driven generative bio-inspired design method according to claim 1, characterized in that: The user in step S05 performs a preliminary evaluation of the design solution, including a subjective evaluation of the design solution.

7. The visual language model driven generative bio-inspired design method according to claim 1, characterized in that: The user in step S06 cooperates with the visual language model to adjust and optimize the prompt words, including the user adjusting and optimizing the prompt words, the visual language model adjusting and optimizing the prompt words, and the user and the visual language model cooperating to adjust and optimize the prompt words; the user in step S06 cooperates with the visual language model to adjust and optimize the prompt words, including the adjustment and optimization of the positive prompt words, the adjustment and optimization of the reverse prompt words, and the adjustment and optimization of the positive prompt words and the reverse prompt words at the same time.

8. The visual language model driven generative bio-inspired design method according to claim 1, characterized in that: The quantitative evaluation of the design scheme based on the visual language model in step S07 includes fine-tuning the visual language model; the quantitative evaluation of the design scheme based on the visual language model in step S07 includes evaluation indicators such as innovation, aesthetics, structural rationality and functional adaptability.

9. The visual language model driven generative bio-inspired design method according to claim 1, characterized in that: This method is widely applicable to innovative designs in fields such as bionic robots, industrial equipment and intelligent manufacturing.