Large-model-based unified representation and reasoning method for multi-modal design information

By constructing a function-behavior-structure graph network, information is extracted from design schemes using a large model and changes in the graph network are compared. This solves the semantic association problem of multimodal design information and enables the optimization of design schemes and timely and accurate feedback.

WO2025251551A1PCT designated stage Publication Date: 2025-12-11ZHEJIANG UNIV

Patent Information

Application Number
PCT/CN2024/134528
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-07
Filing Date
2024-11-26
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing multimodal large models cannot fully understand and utilize the semantic relationships between multimodal design information in product concept design, resulting in a lack of timely and accurate feedback, and an inability to effectively optimize design ideas and processes.

Method used

By constructing a graph network of function, behavior, and structure, structural, functional, and behavioral information is extracted from the current design scheme using a large model. By comparing the changes in the graph network at different times, optimization schemes in each dimension are obtained, and finally, the optimized design scheme is obtained by aggregation.

Benefits of technology

It enables accurate and unified expression and understanding of multimodal information of design schemes, and can optimize design ideas and processes in a timely and accurate manner, providing logically complete reasoning results that are highly relevant to the current design process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024134528_11122025_PF_FP_ABST
    Figure CN2024134528_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a large-model-based unified representation and reasoning method for multi-modal design information. In the present invention, a large model is used to extract corresponding information of function-behavior-structure from a design scheme, so as to construct a function-behavior-structure graph network, thereby achieving a relatively accurate unified representation of multi-modal design information. In the present invention, on the basis of the comparison between graph networks at different moments, the differences in three dimensions, i.e., function-behavior-structure, are obtained, and on the basis of the differences in each dimension, an optimization scheme for the corresponding dimension can be obtained using a large model, so that design thinking and a design process of a designer in each dimension can be obtained by means of reasoning using a large model, a design scheme for each dimension is optimized on the basis of the design thinking and the design process, and corresponding optimization schemes for the three dimensions are then aggregated to obtain an optimized design scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Unified expression and reasoning method for multi-modal design information based on large model TECHNICAL FIELD

[0001] The present application belongs to the field of computer-aided design, and specifically relates to a unified expression and reasoning method for multi-modal design information based on a large model. BACKGROUND

[0002] Product conceptual design is a complex process of generating product schemes according to design tasks, which requires designers to analyze, understand and infer the information contained in the design scheme in depth to capture the key information in abstract and ambiguous user requirements and reason out feasible high-quality design schemes. Traditional computer-aided conceptual design reasoning methods (such as instance reasoning, qualitative reasoning and reasoning using neural networks, etc.) mostly imitate and follow existing information such as design instances and design rules, and fail to truly understand the evolution process of the scheme. These reasoning methods have high requirements for the correlation between auxiliary information and design tasks, and thus are usually only applicable to specific products, lacking general design knowledge and generalization ability.

[0003] The brain-like cognitive model of product conceptual innovative design mainly processes design information presented in a multi-modal manner and forms a semantic cognitive space to understand, express and deduce design information, realizing the knowledge of design information. Existing mainstream computer-aided design or computer-aided conceptual design technologies are developed around a single medium, such as natural language, symbolic logic, two-dimensional sketches, two-dimensional drawings and three-dimensional models. However, human designers process design information and promote the design process in a multi-modal and cross-media manner. In the traditional computer-aided design mode, the processing of design information between human designers and intelligent systems is quite different in form and difficult to integrate. The brain-like cognitive model of conceptual innovative design has two new features: one is to enable computers or intelligent systems to have multi-modal design information expression and processing capability, and the other is to provide a basis for human-machine collaborative design information processing. Multi-modal design information is represented as text, images, digital models and neural networks, and multi-modal information representation and processing is a basic feature of brain-like cognition. Human-machine collaborative design information processing is to use intelligent algorithms inspired by brain-like cognition to enhance the problem definition, design solving and inspiration of intelligent systems and human designers.

[0004] In the human-machine collaborative design process, large language models (LLM) are gradually becoming a powerful assistant for designers. Using LLM for product conceptual design scheme reasoning can effectively solve the problems of requiring auxiliary knowledge and lacking generalization in traditional reasoning methods by utilizing its powerful text understanding ability and massive knowledge reserve. In addition, the development of multi-modal large models provides support for understanding, converting and generating multi-modal data. Using multi-modal large models can integrate multi-modal design information in the design scheme and enhance the comprehensiveness and accuracy of understanding the design scheme.

[0005] However, the processing of multi-modal data by existing multi-modal large models only stops at the alignment of modalities, that is, mapping multi-modal data to a text space that can be understood by the LLM. Using only multi-modal large models to understand the design scheme will lack sufficient mining of the semantic association between multi-modal design information. Therefore, there is still a lack of a product concept design scheme reasoning method that can simultaneously understand multi-modal information of a design scheme and its evolution process. In other words, in the existing concept design reasoning method involving a large model, the feedback provided by the large model is mostly a response to the questions or instructions of the designer, and the design information in the complete design scheme is not fully understood and utilized, and the designer's design ideas and the current design process are not able to provide timely and accurate feedback. SUMMARY

[0006] The present application provides a multi-modal design information unified expression and reasoning method based on a large model, which can convert important information in a design scheme, i.e., function, structure and behavior information and their connections, into a graph network, thereby more accurately realizing multi-modal unified expression, and based on the changes in the graph network at different times under multi-modal unified expression, enabling the large model to accurately and timely understand the designer's ideas and design process, and thereby realizing optimization of the design scheme.

[0007] The present application provides a multi-modal design information unified expression and reasoning method based on a large model, which can convert important information in a design scheme, i.e., function, structure and behavior information and their connections, into a graph network, thereby more accurately realizing multi-modal unified expression, and based on the changes in the graph network at different times under multi-modal unified expression, enabling the large model to accurately and timely understand the designer's ideas and design process, and thereby realizing optimization of the design scheme.

[0008] Based on the graph network at the previous time, the large model extracts new structure nodes and edges between the new structure nodes and other structure nodes from the current design scheme to update the graph network at the previous time to obtain a first graph network, based on the first graph network, the large model extracts new function nodes and edges between the new function nodes and the new structure nodes and / or other structure nodes from the current design scheme to update the first graph network to obtain a second graph network, based on the second graph network, the large model extracts new behavior nodes and edges between the new behavior nodes and the new function nodes and / or other function nodes from the current design scheme to update the second graph network to obtain a graph network at the current time;

[0009] By comparing the graph networks at different times using a mind map, the large model obtains a description of changes in the three dimensions of function, structure and behavior, based on the description of changes in the three dimensions of function, structure and behavior, the large model respectively obtains a plurality of design schemes corresponding to each dimension, based on the set evaluation rules, the large model respectively evaluates the plurality of design schemes of each dimension to obtain optimal design schemes corresponding to function, structure and behavior, and aggregates the optimal design schemes corresponding to function, structure and behavior to obtain an optimized design scheme at the current time.

[0010] Further, the first graph network based on the last time utilizes a large model to extract a new structure node and an edge between the new structure node and other structure nodes from the current design scheme, including:

[0011] The first prompt word is input into the large model to obtain the new structure node, and the first prompt word is constructed by a first instructive sentence, a first output example, a current design scheme and a graph network of the last time. The first instructive sentence makes the large model obtain a first instruction, and the first instruction is to obtain the new structure node. The first output example makes the sentence structure of the result output by the large model be the same as that of the first output example;

[0012] The second prompt word is input into the large model to obtain the edge between the new structure node and other structure nodes, and the second prompt word is constructed by a second instructive sentence, a second output example, a structure node set, a current design scheme and a graph network of the last time. The second instructive sentence makes the large model obtain a second instruction, and the second instruction is to obtain the edge between the new structure node and other structure nodes. The second output example makes the sentence structure of the result output by the large model be the same as that of the second output example. The structure node set includes the new structure node and other structure nodes, and the other structure nodes are structure nodes contained in the graph network of the last time.

[0013] Further, the first graph network based on the last time utilizes a large model to extract a new structure node and an edge between the new structure node and other structure nodes from the current design scheme, including:

[0014] The third prompt word is input into the large model to obtain the new function node, and the third prompt word is constructed by a third instructive sentence, a third output example, a current design scheme, a graph network of the last time and a structure node set. The third instructive sentence makes the large model obtain a third instruction, and the third instruction is to obtain the new function node. The third output example makes the sentence structure of the result output by the large model be the same as that of the third output example. The structure node set includes the new structure node and other structure nodes, and the other structure nodes are structure nodes contained in the graph network of the last time.

[0015] The fourth prompt word is input into the large model to obtain the edge between the new function node and the new structure node and / or other structure nodes, and the fourth prompt word is constructed by a fourth instructive sentence, a fourth output example, a current design scheme, a graph network of the last time and a structure node set. The fourth instructive sentence makes the large model obtain a fourth instruction, and the fourth instruction is to obtain the edge between the new function node and the new structure node and / or other structure nodes. The fourth output example makes the sentence structure of the result output by the large model be the same as that of the fourth output example.

[0016] Further, the second graph network based on the large model extracts a new behavior node from the current design scheme and an edge between the new behavior node and a new function node and / or other function nodes, including:

[0017] inputting a fifth prompt word into the large model to obtain a new behavior node, wherein the fifth prompt word is constructed by a fifth instructive sentence, a fifth output example, a current design scheme, a graph network at a previous time, a structure node set, and a function node set, the fifth instructive sentence makes the large model obtain a fifth instruction, the fifth instruction is to obtain a new behavior node, the fifth output example makes the sentence structure of the result output by the large model be the same as that of the fifth output example, the structure node set includes a new structure node and other structure nodes, the other structure nodes are structure nodes contained in the graph network at the previous time, and the function node set includes a new function node and other function nodes, the other function nodes are function nodes contained in the graph network at the previous time;

[0018] inputting a sixth prompt word into the large model to obtain an edge between the new behavior node and a new function node and / or other function nodes, wherein the sixth prompt word is constructed by a sixth instructive sentence, a sixth output example, a current design scheme, a graph network at a previous time, a structure node set, and a function node set, the sixth instructive sentence makes the large model obtain a sixth instruction, the sixth instruction is to obtain an edge between the new behavior node and a new function node and / or other function nodes, and the sixth output example makes the sentence structure of the result output by the large model be the same as that of the sixth output example.

[0019] Further, a method for constructing a first graph network includes:

[0020] S1, extracting a target product from an initial design scheme by using a large model, and taking the target product as a first structure node;

[0021] S2, extracting a structure node related to the first structure node and an edge between the first structure node and the structure node related to the first structure node from a current design scheme by using the large model based on the extracted first structure node, to obtain a structure node graph network;

[0022] S3, extracting a corresponding function node and an edge between the function node and a structure node of the structure node graph network from the current design scheme by using the large model based on the structure node graph network, and updating the structure node graph network based on the corresponding function node and the edge between the function node and the structure node of the structure node graph network to obtain a graph network containing structure nodes and function nodes;

[0023] S3, based on the graph network comprising the structure node and the function node, utilizing the large model to propose the corresponding behavior node from the current design scheme and the edge of the behavior node and the function node obtained in step S2, updating the graph network based on the corresponding behavior node and the edge of the behavior node and the function node obtained in step S2 to obtain a first graph network.

[0024] Further, the thought graph compares the structure of the graph network at the last moment and the current moment by the large model to obtain the change description of the three dimensions of function, structure and behavior, including:

[0025] The thought graph decomposes the graph network at the last moment and the current moment into three dimensions of function, behavior and structure respectively, and compares the differences of each dimension at different moments by the large model, so as to obtain the change description of the three dimensions of function, structure and behavior.

[0026] Further, the thought graph comprises a controller, a prompt word generator, a parser and an evaluation module.

[0027] The controller is configured to construct an inference state graph and an operation graph based on the graph network data at different moments, wherein the inference state graph comprises a plurality of thought nodes and edges between the thought nodes, the thought nodes comprise four levels of thought nodes, the first level of thought nodes are the change descriptions of the three dimensions of function, structure and behavior, the second level of thought nodes are a plurality of design schemes corresponding to each dimension, the third level of thought nodes are optimal design schemes corresponding to each dimension, and the fourth level of thought nodes are the optimized design scheme at the current moment, the edges between the thought nodes are used to represent the connection relationship between the thought nodes at different levels, and the operation graph is an operation flow constructed by a plurality of operation instructions from the first level of thought nodes to the fourth level of thought nodes based on the inference state graph.

[0028] The prompt word generator is configured to generate a corresponding prompt word based on each operation instruction, and input the prompt word into the large model to generate description information.

[0029] The parser is configured to extract key information from the description information and structure the key information as the pointed thought node.

[0030] The evaluation module is configured to evaluate a plurality of design schemes of each dimension by the large model based on the set evaluation rule to obtain the optimal scheme corresponding to the three dimensions of function, structure and behavior.

[0031] Further, the operation instruction makes the thought node at the current level point to the thought node at the next level, and the operation instruction comprises generation, aggregation, improvement, scoring and selection.

[0032] Further, before inputting the graph network into the large model, the graph network is stored in the form of an adjacency list, and the stored graph network is converted into a natural language description in the form of a string.

[0033] Compared with the prior art, the beneficial effects of the present application are:

[0034] The present application extracts the corresponding information of function-behavior-structure in the design scheme by using a large model, so as to construct a graph network of function-behavior-structure, thereby realizing more accurate unified expression of multi-modal design information.

[0035] Based on the comparison of the graph network of each dimension at different times, the present application obtains the differences in the three dimensions of function-behavior-structure, and uses a large model to obtain the optimization scheme of the corresponding dimension through the differences in each dimension, so as to realize that the design ideas and design processes of the designer under each dimension can be inferred by the large model, and the design scheme of each dimension is optimized based on the design ideas and design processes, and then the optimized design scheme is obtained by aggregating the optimization schemes corresponding to the three dimensions. BRIEF DESCRIPTION OF DRAWINGS

[0036] Fig. 1 is a flowchart of a method for unified expression and reasoning of multi-modal design information based on a large model according to an embodiment of the present application;

[0037] Fig. 2 is a flowchart of converting a current design scheme into a graph network at a current time according to an embodiment of the present application;

[0038] Fig. 3 is a flowchart of obtaining an optimized design scheme according to an embodiment of the present application. DETAILED DESCRIPTION

[0039] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.

[0040] Based on the prior art, the large model cannot accurately obtain unified expression of multi-modal design information, and cannot timely and accurately optimize the design scheme based on the design intent and process. The embodiments of the present application utilize a graph network to more accurately structure the function-behavior-structure (FBS) information in a multi-modal design scheme, and by comparing the graph networks at different times, the large model can more accurately learn the design intent and process, and then reasonably optimize the current design scheme. The specific description of the method for unified expression and reasoning of multi-modal design information based on a large model according to the embodiments of the present application is as follows.

[0041] The generation process of the large model provided by the embodiment of the present application can be uniformly represented as Wherein, p θ represents a pre-trained large model, and represents the parameters of the large model; prompt n represents a prompt word designed for different tasks n, that is, the input of the large model; Y represents the output of the large model; is a parsing function for task n, and processes the output content into storable data R.

[0042] Based on the definitions of the three design concepts of function, behavior and structure in the FBS model, the embodiment of the present application extracts corresponding information in the multi-modal design scheme by using the LLM, constructs a function-behavior-structure graph, and thus realizes the structured representation of the multi-modal design information. By converting the graph network constructed at the last moment into a natural language description and adding it to the prompt word, the LLM is guided to update the nodes and edges on this basis to obtain the graph network at the current moment, so as to construct a discrete dynamic graph through the graph networks at multiple moments, realize the modeling of the deduction process of the multi-modal conceptual design scheme.

[0043] In a specific embodiment, the embodiment needs to structure the long and complex multi-modal design scheme through a graph network, so it is necessary to first extract the key design concepts therein and classify them appropriately, so that the connection between the design concepts is more logical. The embodiment uses the FBS model widely used in the field of product design to extract the three types of key design concepts of function, behavior and structure contained in the multi-modal design scheme.

[0044] The structure provided by the embodiment of the present application includes the target product and its component elements, which are usually actual existing physical entities, and represents what the design object is; the function provided by the embodiment of the present application is the description of the design intention of the target product itself and its components in the design scheme, which is closely related to the user demand, and explains why a certain design object is introduced; the behavior provided by the embodiment of the present application is some attributes brought or expected to be brought by the structure, that is, what the design object can do.

[0045] The embodiment of the present application records a series of graph networks at a series of moments through a graph sequence constructed by discrete graph networks, and the embodiment of the present application models the evolution process of the design scheme by using the discrete dynamic graph. After the designer selects to save the design scheme X t at the current moment t t-1 , new functions, behaviors, structures and the relationships therebetween are extracted from X t , and are added to G t-1 as nodes and edges respectively, so as to obtain a new graph network Gt Thereby, discrete dynamic graphs {G0, G1,..., Gt} are obtained, and the deduction process modeling of the multi-modal conceptual design scheme is realized, wherein t} is realized.

[0046] Specifically, the corresponding prompt word guides the multi-modal large model to extract the design concepts of the newly added functions, behaviors, structures, etc. in the design scheme X t t. The prompt word prompt n needs to contain an instructional sentence instruction n , the design scheme X t at the current time, and the graph network G t-1 at the last time. An example of the instructional sentence is as follows: "extract new component elements as structural nodes from the design scheme at the current time based on the graph network at the last time. The instructional sentence only gives keywords separated by a colon, without specific elaboration". In addition, at each extraction, a task-related output example example is constructed in prompt n for the LLM to learn the context, so that the output of the LLM is more structured. For example, the example can be "electric scooter, handlebar, body". Therefore, prompt n = <instruction n , X t , G t-1 , example, ε>. Wherein ε represents additional information added for a specific task.

[0047] The construction method of the function-behavior-structure graph network at a certain time provided by the embodiment of the application uses a multi-modal large model p θ to extract function, behavior, structure and other design concepts from the multi-modal design scheme X t to form the node set V = F U B U S of the graph. Wherein t represents the time, F, B and S represent the function node, the behavior node and the structure node respectively. The nodes are connected according to the semantic relationship between the design concepts, and the edge set E is constructed. Then the "function-behavior-structure" graph G = (V, E) is constructed. As shown in FIG. 1, the specific steps are as follows:

[0048] (1) The structure, function and behavior information of the current design scheme is extracted by using the large model, and the graph network is constructed based on the extracted structure, function and behavior information to realize unified expression of multi-modal design information. The method for constructing the first graph network provided by the embodiment of the application comprises:

[0049] S1. First, LLM is needed to identify the target product from the multimodal design information of the initial design scheme X1, and use it as the first structural node in the set of structural nodes. Using this as the center of the graph, the set of structural nodes is initialized to S = {s0}. The corresponding prompt word is "prompt". obj = <instruction obj ,X1,G0,example>,instruction obj Extract the target product (object, abbreviated as obj) for the task.

[0050] S2. Based on the first extracted structural node, utilize the large model from the current design scheme X. t Extract the structural nodes associated with the first structural node. And update the set of structure nodes S′=S∪ΔS. Here, i represents the index of the first structure node extracted this time, and its value is equal to the number of existing structure nodes; k is the number of structure nodes extracted this time minus one; the prompt word provided in this embodiment for extracting structural concepts (struct) from large input models is "prompt". struct = <instruction struct ,X t G t-1 Then, LLM is used to mine the new relationships that may appear between the updated structural nodes, that is, the edges between the first structural node and its related structural nodes. And update the edge set E′=E∪ΔE. Where (s m ,s j ) represents the m-th node s m To the j-th node s j The directed path s between m →s j s m ,s j ∈S′. The prompt word provided in this embodiment for extracting the structural edge concept from a large input model is "prompt". edge = <instruction edge ,X t G t-1 Example, ε>,ε=ΔS. The structural node graph network is obtained by connecting the structural nodes and their corresponding edges.

[0051] S3, each structural node s based on the structural node graph network. uextracts the relevant design requirements and target functions of the current design scheme from the LLM as a function node wherein u represents the index of the structure node; the value of w is the number of function nodes related to the structure node s u minus one; the prompt provided by the embodiment for inputting the task extraction function concept (function, abbreviated as func) of the large model is prompt func <instruction func ,X t ,G t-1 ,example,ε> and ε = ΔS. Update the function node set F' = F U ΔF, and update the edges of the structure nodes of the function node and structure node graph network, and the updated edge set E' = E U ΔE, wherein ΔE = {(s u ,f uw )|f uw ∈ΔF}, and the obtained function node and function node and structure node edges are added to the structure node graph network to obtain a binary graph network of structure nodes and function nodes.

[0052] S4, based on each "structure-function" binary tuple (s u ,f ug ), s u ∈S', f ug ∈F', the LLM extracts the structure s u from the current design scheme to realize the function f ug and the related attributes b ugh as a behavior node wherein u represents the index of the structure node; g represents the index of the function node connected to the structure node s u ; the value of h is the number of behavior nodes related to the function node f ug extracted this time minus one; the prompt provided by the embodiment for inputting the task extraction behavior concept (behavior, abbreviated as behav) of the large model is prompt behav <instruction behav ,X t ,G t-1 ,example,ε> and ε = ΔS + ΔF. Update the behavior node set B' = B U ΔB, and update the edge set E' = E U ΔE, wherein ΔE = {(f ug ,b ugh )|b ughThe first graph network is obtained by adding the obtained corresponding behavior node and the edge between the behavior node and the function node obtained in step S2 to the structure node graph network.

[0053] The embodiment of the present application finds new structures, functions, behaviors and edges between them from the current design scheme based on the graph network of the previous moment by using a large model, thereby constructing the current moment graph network corresponding to the current design scheme, as shown in FIG. 2, in each graph network in the dynamic graph provided by the embodiment of the present application, the graph network of the previous moment G t-1 The graph network G of the current moment is obtained t , comprising:

[0054] The new structure node and the edge between the new structure node and other structure nodes are extracted from the current design scheme based on the graph network of the previous moment by using a large model, and the graph network of the previous moment is updated based on the extracted new structure node and the edge between the new structure node and other structure nodes to obtain the first graph network.

[0055] In one embodiment, the specific steps of the embodiment of the present application for extracting the new structure node and the edge between the new structure node and other structure nodes from the current design scheme based on the graph network of the previous moment by using a large model are as follows:

[0056] In this embodiment, the first prompt word is input into the large model to obtain the new structure node, and the first prompt word is constructed by a first instructive sentence, a first output example, a current design scheme and a graph network of the previous moment. The first instructive sentence makes the large model obtain a first instruction, and the first instruction is to obtain a new structure node. The first output example makes the sentence structure of the result output by the large model be the same as that of the first output example.

[0057] In this embodiment, the second prompt word is input into the large model to obtain the edge between the new structure node and other structure nodes, and the second prompt word is constructed by a second instructive sentence, a second output example, a structure node set, a current design scheme and a graph network of the previous moment. The second instructive sentence makes the large model obtain a second instruction, and the second instruction is to obtain the edge between the new structure node and other structure nodes. The second output example makes the sentence structure of the result output by the large model be the same as that of the second output example. The structure node set includes the new structure node and other structure nodes, and the other structure nodes are structure nodes contained in the graph network of the previous moment. In one embodiment, the structure node set is stored in the form of an adjacency list.

[0058] The embodiment of the present application extracts a new function node and an edge between the new function node and a new structure node and / or other structure nodes from the current design scheme based on the first graph network using a large model, and updates the first graph network based on the extracted new function node and the edge between the new function node and the new structure node and / or other structure nodes to obtain a second graph network.

[0059] In an embodiment, the embodiment of the present application provides specific steps for extracting a new function node and an edge between the new function node and a new structure node and / or other structure nodes from the current design scheme based on the first graph network using a large model, which are as follows:

[0060] In this embodiment, a third prompt word is input into the large model to obtain a new function node, wherein the third prompt word is constructed by a third instructive sentence, a third output example, a current design scheme, a graph network at a previous time and a structure node set. The third instructive sentence makes the large model obtain a third instruction, which is to obtain a new function node. The third output example makes the sentence structure of the result output by the large model be the same as that of the third output example. The structure node set includes a new structure node and other structure nodes, and the other structure nodes are structure nodes contained in the graph network at the previous time. In an embodiment, the structure node set is stored in the form of an adjacency list.

[0061] In this embodiment, a fourth prompt word is input into the large model to obtain a new function node and an edge between the new function node and a new structure node and / or other structure nodes, wherein the fourth prompt word is constructed by a fourth instructive sentence, a fourth output example, a structure node set, a current design scheme, a graph network at a previous time and a structure node set. The fourth instructive sentence makes the large model obtain a fourth instruction, which is to obtain a new function node and an edge between the new function node and a new structure node and / or other structure nodes. The fourth output example makes the sentence structure of the result output by the large model be the same as that of the fourth output example.

[0062] The embodiment of the present application extracts a new behavior node and an edge between the new behavior node and a new function node and / or other function nodes from the current design scheme based on the second graph network using a large model, and updates the second graph network based on the extracted new behavior node and the edge between the new behavior node and the new function node and / or other function nodes to obtain a graph network at a current time.

[0063] In an embodiment, the embodiment of the present application provides specific steps for extracting a new behavior node and an edge between the new behavior node and a new function node and / or other function nodes from the current design scheme based on the second graph network using a large model, which are as follows:

[0064] The fifth prompt word is input into the large model to obtain a new behavior node, wherein the fifth prompt word is constructed by a fifth instructing sentence, a fifth output example, a current design scheme, a graph network at a previous moment, a structure node set and a function node set, the fifth instructing sentence makes the large model obtain a fifth instruction, the fifth instruction is to obtain a new behavior node, the fifth output example makes the sentence structure of the result output by the large model be the same as the sentence structure of the fifth output example, the structure node set includes a new structure node and other structure nodes, the other structure nodes are structure nodes contained in the graph network at the previous moment, the function node set includes a new function node and other function nodes, and the other function nodes are function nodes contained in the graph network at the previous moment.

[0065] The sixth prompt word is input into the large model to obtain an edge between the new behavior node and the new function node and / or other function nodes, wherein the sixth prompt word is constructed by a sixth instructing sentence, a sixth output example, a structure node set, a current design scheme, a graph network at a previous moment, a structure node set and a function node set, the sixth instructing sentence makes the large model obtain a sixth instruction, the sixth instruction is to obtain an edge between the new behavior node and the new function node and / or other function nodes, and the sixth output example makes the sentence structure of the result output by the large model be the same as the sentence structure of the sixth output example.

[0066] The graph network data of function-behavior-structure is stored in the form of an adjacency list in the embodiment of the application, and specifically, two classes of node classes and graph classes of the graph network are defined.

[0067] The node class Vertex defined in the embodiment of the application is for each node in the graph, and the node includes id, type, value and neighbors, the id is an integer variable, represents the serial number of the node, the type is a string variable, represents the type of the node, includes three values of “f”, “b” and “s”, respectively corresponding to function, behavior and structure, the value is a string variable, represents the semantics of the node, and the neighbors are list variables, store the serial numbers of all adjacent nodes pointed by the node.

[0068] The graph defined by the embodiment of the application is the overall architecture of the graph network, including graph_list, time, addVertex, addEdge and is_Vertex_in_Graph, wherein graph_list is a list type variable, storing instances of all nodes in the graph, time is a string variable, the time of establishing the graph, addVertex is a function for adding a new node to graph_list, addEdge is a function for updating edges by updating the neighbors list of a specific node in graph_list, and is_Vertex_in_Graph is a function for judging whether a node exists in the graph.

[0069] The embodiment of the application provides that when the graph needs to be updated, a Vertex instance is created for a new node, and the addVertex function is called to add it to the graph. The addition of edges is realized by calling the addEdge function, and before adding, the is_Vertex_in_Graph function needs to be called to judge whether the two nodes connected by the edge are already in the graph. If not, the nodes need to be added first.

[0070] (2) As shown in FIG. 3, the design process is perceived and the current design scheme is reasoned based on the structure of the graph network at different times using the mind map: the mind map compares the structure of the graph network at different times through a large model to obtain the change description of the three dimensions of function, structure and behavior, as shown in the first layer of FIG. 3, that is, the connection relationship between the nodes at the current time (structure node, function node or behavior node, dark color dot in FIG. 3) and the nodes at the last time (structure node, function node or behavior node, light color dot in FIG. 3) based on Δt time is obtained through the large model. The change description of the three dimensions of function, structure and behavior, based on the change description of the three dimensions of function, structure and behavior, a plurality of design schemes corresponding to each dimension are obtained by using a large model, and based on the set evaluation rule, a plurality of design schemes of each dimension are evaluated by using a large model to obtain the optimal design scheme corresponding to function, structure and behavior. The optimal design schemes corresponding to function, structure and behavior are aggregated to obtain the optimized design scheme at the current time.

[0071] The embodiment of the application adopts the prompt word fine-tuning method of the mind map, decomposes the graph network containing "function-behavior-structure" into three dimensions of function, behavior and structure, guides the LLM to understand the design ideas of the designer and the current design process on different dimensions by comparing the differences between graphs at different times. On this basis, a plurality of creative stimuli, i.e. a plurality of design schemes, are generated for different dimensions, and the LLM or the designer scores and compares them to select the optimal scheme under each dimension. Through the aggregation and optimization of these schemes, a logically complete and highly relevant reasoning result to the current design process is obtained.

[0072] The thought graph provided by the embodiment of the present application is a prompt word fine-tuning method, which models the reasoning process of the LLM through a graph structure, realizes the decomposition of the task and the divergence and aggregation of thinking, and thus enhances the reasoning ability of the LLM for complex logic. The basic framework of the thought graph includes a controller, a prompt word generator, a parser and an evaluation model.

[0073] As shown in FIG. 3, the controller provided by the embodiment of the present application includes a reasoning state graph and an operation graph. The reasoning state graph includes a plurality of thought nodes and edges between the thought nodes. The thought nodes include four levels of thought nodes. The thought nodes of the first level are changes in three dimensions of function, structure and behavior. The thought nodes of this level are determined by comparing the graph networks at different times obtained in step (1). The thought nodes of the second level are a plurality of design schemes corresponding to each dimension obtained by reasoning and generating operation instructions based on the change description of each dimension. The design schemes corresponding to the structure dimension contain structure concepts, the design schemes corresponding to the function dimension contain function concepts, and the design schemes corresponding to the behavior dimension contain behavior concepts. The thought nodes of the third level are optimal design schemes selected from the plurality of design schemes of each dimension through scoring and selection operation instructions. The thought nodes of the fourth level are optimized design schemes at the current time obtained by aggregating operation instructions of the optimal design schemes of each dimension.

[0074] The edges between the thought nodes provided by the embodiment of the present application are used to represent the connection relationship between the thought nodes of different levels.

[0075] The operation graph provided by the embodiment of the present application is an operation flow constructed based on a plurality of operation instructions from the thought nodes of the first level to the thought nodes of the fourth level obtained based on the reasoning state graph.

[0076] The prompt word generator provided by the embodiment of the present application is used to generate corresponding prompt words based on each operation instruction, and input the prompt words into a large model to generate description information.

[0077] In an embodiment, the prompt word generator provided by the embodiment of the present application selects a pre-set prompt word framework corresponding to the operation (generation, aggregation or improvement) to be performed according to the operation to be performed in the current thought state node flow control information and the reasoning stage, embeds a variable containing specific input information, generates a corresponding prompt word as an input of the LLM, and guides the LLM to generate a corresponding feedback.

[0078] The parser provided by the embodiment of the present application is used to extract key information from the description information and structure the key information as a pointed thought node. The key information is content information that can obtain the pointed thought node.

[0079] It can be understood that the scoring stage provided by the embodiment of the present application returns a score and a reason for the score, and the next stage only needs the score information without the reason, and the resolver extracts the score as the thought node from the content returned by the LLM.

[0080] The evaluation module provided by the embodiment of the present application is used for evaluating a plurality of design schemes in each dimension by a large model based on a set evaluation rule to obtain optimal schemes corresponding to the three dimensions of function, structure and behavior.

[0081] In an embodiment, the evaluation module provided by the embodiment of the present application scores or verifies the correctness of the design scheme under each dimension represented by the node in the reasoning state graph. The evaluation can be completed by the LLM, or by the designer, or the evaluation results of both are considered to enhance the rationality of the evaluation.

[0082] In order to make the score of the scheme generated in the reasoning intermediate process by the LLM more reasonable and more interpretable, and to ensure the consistency of the evaluation criteria for different schemes, the embodiment of the present application provides the following evaluation dimensions for the product conceptual design scheme of the LLM: 1. Timeliness: whether the scheme content is highly related to the current design process. 2. Innovation: whether the product design provides a new solution or uses innovative technology, and whether it is significantly distinguished from existing products. 3. User demand satisfaction: whether the design scheme solves the actual problem of the user and whether it meets the user's demand. 4. Functionality: whether the function of the product is complete, whether it can effectively perform its predetermined function, and whether the function corresponds to the user's demand. 5. Ease of use: whether the product is easy to use, and how the user experiences in the use process, including the friendliness of the interface and the intuitiveness of the operation. 6. Technical feasibility: whether the technology relied on by the design scheme is mature and can be actually manufactured or implemented. 7. Cost-effectiveness: evaluate the product design from the cost perspective, including production cost, maintenance cost, etc. 8. Aesthetic design: whether the appearance, color, shape, etc. of the product attract the target user group and whether it meets the aesthetic trend.

[0083] In an embodiment, the reasoning of the design scheme by the LLM of the embodiment of the present application is performed on the extracted graph network containing "function-behavior-structure". The embodiment converts the structured stored graph network into natural language input to the LLM, and the specific steps of preprocessing the structured stored graph network data into natural language description are as follows:

[0084] In each time of completing the construction of the "function-behavior-structure" graph, four strings are initialized: str_s="structure", str_f="function", str_b="behavior", and str_e="relationship".

[0085] Then, iterate through each node element `node` in the `graph_list` instance of the graph network, concatenating its semantic value and index as strings after the corresponding category based on the node type, i.e., `str_{node.type} += "{node.value}({node.id})"`. The elements in the `neighbors` member variable of the node represent the index of the neighboring nodes pointed to by that node, denoted as `num`. Iterate through the elements in `node.neighbors`, concatenating the tuple `(node.id, num)` as strings after the "relation", i.e., `str_e += "({node.id}, {num})"`. Finally, concatenating these four strings yields the "functional-behavioral-structural" graph G. t Natural Language Description D t = str_f + str_b + str_s + str_e.

[0086] In a specific embodiment of the present invention, the natural language descriptions of the k nearest images {D} are selected. t-k+1 ,...,D t The initial node of the reasoning state diagram in the mind map framework, as the core idea, begins the reasoning process according to the operation diagram.

[0087] In a specific embodiment of the present invention, operation instructions are used to make a thought node at one level point to a thought node at the next level. The operation instructions include generating, aggregating, improving, scoring, and selecting.

[0088] This embodiment provides a reasoning process that generates multiple ideas from a single idea. It utilizes LLM to decompose large tasks into multiple subtasks, or to generate multiple solutions for a single problem. The generation operation is represented on the reasoning state graph as the splitting of an idea node v into multiple new idea nodes, where M is the number of new idea nodes. The set of new idea nodes and the edges connecting the new nodes to other nodes are as follows:

[0089] The aggregation method provided in this embodiment is a reasoning process that integrates multiple ideas into a single idea. It utilizes LLM to fuse and enhance the advantages of multiple ideas while eliminating conflicts between different ideas. In the reasoning state graph, the aggregation operation is represented by multiple nodes pointing to a new idea node v. + The new set of thought nodes and the edges connecting the new node to other nodes are: ΔV T ={v +}, ΔE T ={(v1,v + ),...,(v M ,v + )}.

[0090] The perfect provided by the embodiment is the process of re-reasoning the current thought by LLM and modifying its content. The perfect operation is embodied as the addition of a loop from the thought node v to itself on the reasoning state diagram, that is ΔE T ={(v,v)}.

[0091] The score provided by the embodiment is to evaluate and score the thought nodes in the reasoning state diagram. LLM can score according to relevant indicators, or designers can score.

[0092] The selection provided by the embodiment is to rank the specified thought nodes according to their scores, and keep the top h nodes with the highest scores.

[0093] In order to make LLM deeply understand the design ideas of designers, and generate logical and highly relevant reasoning results for the current design process, the specific embodiments of the present application have developed the following operation process:

[0094] According to the natural language description input={D t-k+1 ,...,D t} of the "function-behavior-structure" graph network at the last k moments, the description of the changes of the design scheme in function, behavior and structure respectively is generated The prompt word needs to guide LLM to analyze the changes of the current stage dynamic graph in structure, function and behavior respectively, and the design problems that the designer may pay attention to in structure, function and behavior respectively, and contains the natural language description of the last k moments "function-behavior-structure" graph and output examples. That is prompt generate0 =<instruction generate0 ,input,example>.

[0095] According to the changes of the above three aspects, the design scheme is reasoned, and three possible schemes are generated for each dimension z The prompt word needs to guide LLM to give further deduction direction or solution to the existing problem of the design scheme according to the analysis of the last reasoning stage, and contains the description of the changes of the design scheme in function, behavior and structure respectively and output examples. That is prompt generate1 =<instruction generate1 ,L z ,example>,z∈{f,b,s}.

[0096] The three schemes for each dimension are scored. LLM gives the score according to the given indicators The prompt word needs to inform the LLM scoring criteria (such as ease of use, feasibility, etc.) and the scoring interval, and contains the generated schemes and output examples respectively for function, behavior, and structure. That is, prompt score = <instruction score , Q, example>, Q ∈ {Q z1 , Q z2 , Q z3 , z ∈ {f, b, s}. The designer gives a subjective score Score2. Then the final score of the scheme Score = α·Score1 + β·Score2, where α and β represent the weights of LLM scoring and designer scoring respectively.

[0097] The specific embodiment of the application retains the highest score of each dimension Q z = max Score (Q z1 , Q z2 , Q z3 ).

[0098] The optimal design scheme generated for function, behavior, and structure changes is combined to obtain the final reasoning result The optimization scheme at the current time. The prompt word needs to guide the LLM to combine the three product concept design schemes, eliminate the conflicting and repetitive parts, and generate the final reasoning result. The prompt word input to the LLM is prompt aggregate Contains the generated schemes and output examples respectively for function, behavior, and structure. That is, prompt aggregate = <instruction aggregate , Q, example>, Q = {Q f , Q b , Q s}.

[0099] After obtaining the optimization scheme at the current time in the specific embodiment of the application, the optimization scheme is converted into the corresponding graph network G t+1 by step (1), and the graph network G t+1 is used for the construction of the next stage graph network and the reasoning of the design scheme.

Claims

1. A large model-based multi-modal design information unified expression and reasoning method, characterized in that, Comprise: Based on the graph network of the last moment, a large model is used to extract new structure nodes and edges between new structure nodes and other structure nodes from the current design scheme to update the graph network of the last moment to obtain a first graph network. Based on the first graph network, a large model is used to extract new function nodes and edges between new function nodes and new structure nodes and / or other structure nodes from the current design scheme to update the first graph network to obtain a second graph network. Based on the second graph network, a large model is used to extract new behavior nodes and edges between new behavior nodes and new function nodes and / or other function nodes from the current design scheme to update the second graph network to obtain the graph network of the current moment; By comparing the graph networks of different moments through a large model using a mind map, a description of changes in three dimensions of function, structure and behavior is obtained. Based on the description of changes in three dimensions of function, structure and behavior, a large model is used to obtain a plurality of design schemes corresponding to each dimension. Based on the set evaluation rules, a large model is used to evaluate a plurality of design schemes of each dimension to obtain optimal design schemes corresponding to function, structure and behavior, respectively. The optimal design schemes corresponding to function, structure and behavior are aggregated to obtain an optimized design scheme of the current moment.

2. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, The large model is used to extract new structure nodes and edges between new structure nodes and other structure nodes from the current design scheme based on the graph network of the last moment, comprising: The first prompt word is input into the large model to obtain the new structure node, and the first prompt word is constructed by the first instructive sentence, the first output example, the current design scheme and the graph network of the last moment. The first instructive sentence makes the large model obtain a first instruction, and the first instruction is to obtain a new structure node. The first output example makes the sentence structure of the result output by the large model the same as that of the first output example; The second prompt word is input into the large model to obtain the new structure node and the edge between the new structure node and the other structure node, and the second prompt word is constructed by the second instructive sentence, the second output example, the structure node set, the current design scheme and the graph network of the last moment. The second instructive sentence makes the large model obtain a second instruction, and the second instruction is to obtain an edge between the new structure node and the other structure node. The second output example makes the sentence structure of the result output by the large model the same as that of the second output example. The structure node set includes the new structure node and the other structure node, and the other structure node is a structure node contained in the graph network of the last moment.

3. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, The large model is used to extract new function nodes and edges between new function nodes and new structure nodes and / or other structure nodes from the current design scheme based on the first graph network, comprising: inputting a third prompt word into the large model to obtain a new function node, wherein the third prompt word is constructed by a third instructive sentence, a third output example, a current design scheme, a graph network at a previous moment, and a structure node set, the third instructive sentence is used to make the large model obtain a third instruction, the third instruction is to obtain the new function node, the third output example is used to make a sentence structure of a result output by the large model be same as a sentence structure of the third output example, and the structure node set includes a new structure node and other structure nodes, the other structure nodes are structure nodes included in the graph network at the previous moment; inputting a fourth prompt word into the large model to obtain a new function node and an edge between the new function node and a new structure node and / or other structure nodes, wherein the fourth prompt word is constructed by a fourth instructive sentence, a fourth output example, a current design scheme, a graph network at a previous moment, and a structure node set, the fourth instructive sentence is used to make the large model obtain a fourth instruction, the fourth instruction is to obtain the new function node and the edge between the new function node and the new structure node and / or the other structure nodes, and the fourth output example is used to make a sentence structure of a result output by the large model be same as a sentence structure of the fourth output example.

4. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, extracting, by the large model, a new behavior node and an edge between the new behavior node and a new function node and / or other function nodes from the current design scheme based on the second graph network, including: inputting a fifth prompt word into the large model to obtain the new behavior node, wherein the fifth prompt word is constructed by a fifth instructive sentence, a fifth output example, a current design scheme, a graph network at a previous moment, a structure node set, and a function node set, the fifth instructive sentence is used to make the large model obtain a fifth instruction, the fifth instruction is to obtain the new behavior node, the fifth output example is used to make a sentence structure of a result output by the large model be same as a sentence structure of the fifth output example, the structure node set includes a new structure node and other structure nodes, the other structure nodes are structure nodes included in the graph network at the previous moment, and the function node set includes a new function node and other function nodes, the other function nodes are function nodes included in the graph network at the previous moment; inputting a sixth prompt word into the large model to obtain the new behavior node and the edge between the new behavior node and the new function node and / or the other function nodes, wherein the sixth prompt word is constructed by a sixth instructive sentence, a sixth output example, a current design scheme, a graph network at a previous moment, a structure node set, and a function node set, the sixth instructive sentence is used to make the large model obtain a sixth instruction, the sixth instruction is to obtain the new behavior node and the edge between the new behavior node and the new function node and / or the other function nodes, and the sixth output example is used to make a sentence structure of a result output by the large model be same as a sentence structure of the sixth output example.

5. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, a method for constructing a first graph network, including: S1, extracting a target product from an initial design scheme by using a large model, and taking the target product as a first structure node; S2, extracting structure nodes related to the first structure node and edges between the first structure node and the structure nodes related thereto from the current design scheme based on the extracted first structure node using a large model, thereby obtaining a structure node graph network; S3, extracting corresponding function nodes and edges between the function nodes and the structure nodes of the structure node graph network from the current design scheme based on the structure node graph network using a large model, updating the structure node graph network based on the corresponding function nodes and the edges between the function nodes and the structure nodes of the structure node graph network to obtain a graph network containing structure nodes and function nodes; S3, extracting corresponding behavior nodes and edges between the behavior nodes and the function nodes obtained in step S2 from the current design scheme based on the graph network containing structure nodes and function nodes using a large model, updating the graph network based on the corresponding behavior nodes and the edges between the behavior nodes and the function nodes obtained in step S2 to obtain a first graph network.

6. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, The thought graph compares the structures of the graph networks at the previous moment and the current moment using a large model to obtain the change descriptions of the three dimensions of function, structure and behavior, including: The thought graph decomposes the graph networks at the previous moment and the current moment into three dimensions of function, behavior and structure, respectively, and compares the differences of each dimension at different moments using a large model, so as to obtain the change descriptions of the three dimensions of function, structure and behavior.

7. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, The thought graph includes a controller, a prompt word generator, a parser and an evaluation module; The controller is configured to construct an inference state graph and an operation graph based on the graph network data at different moments, wherein the inference state graph includes a plurality of thought nodes and edges between the thought nodes, the thought nodes include four levels of thought nodes, the first level of thought nodes are change descriptions of the three dimensions of function, structure and behavior, the second level of thought nodes are a plurality of design schemes corresponding to each dimension, the third level of thought nodes are optimal design schemes corresponding to each dimension, and the fourth level of thought nodes are optimized design schemes at the current moment, the edges between the thought nodes are used to represent the connection relationship between the thought nodes at different levels, and the operation graph is an operation flow constructed by a plurality of operation instructions from the first level of thought nodes to the fourth level of thought nodes based on the inference state graph. The prompt word generator is configured to generate corresponding prompt words based on each operation instruction, and input the prompt words into a large model to generate description information; The parser is configured to extract key information from the description information and structure the key information as pointed thought nodes; The evaluation module is configured to evaluate a plurality of design schemes of each dimension based on a set evaluation rule through a large model to obtain optimal schemes corresponding to the three dimensions of function, structure and behavior.

8. The large model-based multi-modal design information unified representation and reasoning method according to claim 7, characterized in that, The operation instructions include generation, aggregation, improvement, scoring and selection.

9. The large model-based multi-modal design information unified representation and reasoning method according to claim 1, characterized in that, Before inputting the graph network into the large model, the graph network is stored in the form of an adjacency list, and the stored graph network is converted into a natural language description in the form of a string.

Citation Information

Patent Citations

  • Multi-agent information fusion method and device, electronic equipment and readable storage medium

    CN114139637A

  • Man-machine collaborative creation method and system supporting combined creativity and electronic equipment

    CN117932048A

  • Multi-modal design information unified expression and reasoning method based on large model

    CN118313285A

  • System and method of facilitating human interactions with products and services over a network

    US11908476B1

  • Multimodal procedural guidance content creation and conversion methods and systems

    US20230343044A1

Cited By

  • A method and system for detecting mode transition requirement conflicts in onboard software based on LLMs

    CN122433879A