A paper folding prediction method and system based on multimodal graph learning
Through the multimodal graph learning method, combined with the graph convolution neural network and the graph attention network, the multimodal features of origami structure are extracted and fused, and the existing origami prediction methods are solved, and efficient and high-precision origami prediction is achieved.
Patent Information
- Application Number
- CN202510213165.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing origami prediction methods have problems such as poor universality, low prediction efficiency and low accuracy of prediction results, especially when facing large-scale and high-complex prediction tasks.
Using a multi-modal graph learning method, the graph topology structure and two-dimensional crease images are extracted through graph convolution neural network, and feature learning and fusion are performed in combination with graph attention network, and the origami prediction results are output using a multi-layer perceptron.
It realizes efficient and high-precision prediction of origami results, improves the universality and computing efficiency of the prediction model, and can better adapt to the challenges of various origami scenarios.
Smart Images

Figure CN119723547B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an origami prediction method and system based on multimodal graph learning. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Origami structure is a new type of structural system that can be folded on a two-dimensional plane according to preset folds (mountain folds, valley folds, etc.) without cutting or stretching. Origami structure has the characteristics of variable shape, strong design, rigidity and superior folding and unfolding ratio, and has good performance in device structure; therefore, origami structure is widely used in aerospace, medical devices, robotics and other fields. For example, a satellite solar panel has a diameter of only 2.5m when folded, but can reach a diameter of 27m when fully unfolded. It has a large folding and unfolding ratio and can use a smaller storage space to obtain a larger practical area.
[0004] The core of origami prediction is to simulate and predict the folding process and final result under the constraints of folding rules based on the initial conditions of origami, such as the size, shape and crease distribution of the folded material. By calculating and inferring the three-dimensional folding shape through the two-dimensional crease map, it can not only achieve the result prediction of artistic origami, but also give full play to the advantages of origami structure in the industrial field.
[0005] The inventors have found that the existing methods for predicting the results of origami structures have some technical problems, such as:
[0006] (1) The common methods for achieving origami results currently rely mainly on manual experience and limited simulation deduction. For example, the designer first designs the creases, and then manually or mechanically implements them according to the preset creases. However, this method has the disadvantages of long construction period and large amount of engineering. In addition, when designing more precise folding structure devices, the precision of the preparation equipment is extremely high. Therefore, it is difficult to put it into practical application on a large scale.
[0007] (2) The design methods of traditional origami structures mainly rely on existing origami rules such as the Huzita-Hatori axiom, Maekawa conditions, Maekawa's Theorem, and 2-Colorable conditions. These theories have clear constraints on the geometric transformation of the folding process, but the calculation process involved is very complicated, which leads to low calculation efficiency.
[0008] (3) Existing simulation prediction software (such as Origami Simulator) that can be used to deduce and calculate the folding process of origami structures has accelerated the prediction speed to a certain extent. However, this type of simulation software needs to be recalculated every time a prediction is made, does not have the ability to generalize learning, and has great limitations when facing complex origami structures. At the same time, more and more studies want to explore new origami prediction methods through deep learning technology. However, the current models based on deep learning technology are difficult to capture the geometric features and shape features of origami structures at the same time when applied to the field of origami result prediction. Therefore, the prediction results are not accurate enough in terms of morphology and geometric structure, resulting in unsatisfactory prediction results.
[0009] In summary, existing origami prediction methods generally have problems such as poor universality, low prediction efficiency and low prediction accuracy when facing large-scale and high-complexity prediction tasks. Summary of the invention
[0010] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides an origami prediction method and system based on multimodal graph learning, which can achieve efficient and high-precision prediction of origami results and improve the universality of the prediction model.
[0011] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0012] A first aspect of the present invention provides an origami prediction method based on multimodal graph learning.
[0013] A paper-folding prediction method based on multimodal graph learning, comprising:
[0014] Obtain folding design parameters and initial crease images;
[0015] generating a graph topology structure and a two-dimensional crease image according to the folding design parameters;
[0016] A graph convolutional neural network is used to extract features from the graph topology and the two-dimensional fold image to obtain a first geometric feature; the initial fold image and the first geometric feature are sent to a graph attention network for feature learning to obtain a second geometric feature and an image feature;
[0017] The second geometric feature and the image feature are fused, and the fused features are analyzed using a multi-layer perceptron to output an origami prediction result.
[0018] Furthermore, the folding design parameters include vertex coordinates, geometric constraints, crease edge sets and crease types.
[0019] Furthermore, vertex coordinates are used as node features, and geometric constraints and crease types are used as edge features.
[0020] Furthermore, a graph convolutional neural network is used to extract features of the graph topology structure and the two-dimensional fold image, including: setting the convolution step size, using the graph convolutional neural network to generate a subgraph starting from the current node of the graph topology structure with the step size as the distance, transmitting information between the node features and edge features of the subgraph, and aggregating the feature information reflected by the two-dimensional fold image to obtain the first geometric feature.
[0021] Furthermore, the graph attention network performs global weight calculation and feature learning on the initial crease image and the node features and edge features in the first geometric features to obtain the second geometric features and image features.
[0022] Furthermore, before using the graph convolutional neural network and the graph attention network for origami prediction, the graph convolutional neural network and the graph attention network need to be trained; to ensure the accuracy of the model training, the data used for training is augmented to increase the amount of data used for training.
[0023] Furthermore, the graph convolutional neural network and the graph attention network are trained, including: defining a fusion loss function according to the prediction task, comparing the prediction results of the fused features with the true labels and calculating the loss values; then, back-propagating the loss values to the graph convolutional neural network and the graph attention network based on the back-propagation algorithm; and iteratively updating the parameters in the graph convolutional neural network and the graph attention network based on the gradient descent algorithm.
[0024] A second aspect of the present invention provides an origami prediction system based on multimodal graph learning.
[0025] A paper-folding prediction system based on multimodal graph learning, comprising:
[0026] The data acquisition module is configured to: acquire folding design parameters and initial crease images;
[0027] An object generation module is configured to: generate a graph topology structure and a two-dimensional crease image according to the folding design parameters;
[0028] The feature extraction module is configured to: use a graph convolutional neural network to extract features from the graph topology and the two-dimensional fold image to obtain a first geometric feature; send the initial fold image and the first geometric feature to a graph attention network for feature learning to obtain a second geometric feature and an image feature;
[0029] The origami prediction module is configured to: perform feature fusion on the second geometric feature and the image feature, analyze the fused features using a multi-layer perceptron and output an origami prediction result.
[0030] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of an origami prediction method based on multimodal graph learning as described in the first aspect of the present invention.
[0031] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of an origami prediction method based on multimodal graph learning as described in the first aspect of the present invention are implemented.
[0032] One or more of the above technical solutions have the following beneficial effects:
[0033] (1) The present invention uses a graph convolutional neural network to extract features from graph topology and two-dimensional fold images to obtain geometric features; the graph attention network can also learn image features while further learning geometric features, and finally fuse the learned image features and geometric features. Therefore, the present invention can accurately fuse multimodal features and meet the constraints in the origami law, and show good adaptability in different types of origami scenarios; the network architecture design enables it to better cope with the challenges of various origami scenarios and improves the universality of the prediction model.
[0034] (2) The present invention combines graph convolutional neural networks and graph attention networks, which can effectively reduce the amount of computation while retaining key information; at the same time, the model has excellent generalization ability, which greatly reduces the computer's computing power consumption when facing similar structures, thereby improving computational efficiency.
[0035] (3) The present invention not only extracts geometric structure features, but also analyzes image features, which can ensure that the prediction results have high accuracy in both shape and specific geometric structure. At the same time, the present invention performs augmentation operations on the data used for training to increase the amount of data used for training, so that the trained model has better prediction capabilities. Therefore, the present invention can achieve high-precision prediction of origami results.
[0036] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0038] Figure 1 This is a flowchart of an origami prediction method based on multimodal graph learning in Example 1 of the present invention.
[0039] Figure 2 This is a data relationship diagram of an origami prediction method based on multimodal graph learning in Example 1 of the present invention.
[0040] Figure 3 This is a schematic diagram of inputting data to the model in Embodiment 1 of the present invention.
[0041] Figure 4 Schematic diagram of the information transmission process of the graph convolutional neural network in Example 1 of the present invention.
[0042] Figure 5 Schematic diagram of the information transmission process of the attention network in Example 1 of the present invention.
[0043] Figure 6 Schematic diagram of the feature fusion process in Embodiment 1 of the present invention.
[0044] Figure 7 This is an example diagram of a two-dimensional crease design diagram in Example 1 of the present invention.
[0045] Figure 8 It is a schematic diagram of crease result prediction corresponding to the two-dimensional crease design diagram in the first embodiment of the present invention. DETAILED DESCRIPTION
[0046] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0047] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0048] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0049] Embodiment 1
[0050] This embodiment discloses an origami prediction method based on multimodal graph learning.
[0051] like Figure 1 , Figure 2As shown, an origami prediction method based on multimodal graph learning can be applied to predict the two-dimensional folding structure of complex origami models, such as predicting the folding pattern of a paper crane. By inputting the folding design parameters of the paper crane, such as vertex coordinates, crease types, geometric constraints, etc., the model can generate its corresponding graph topology and use the graph convolutional neural network (GCN) to extract features of the crease image and geometric information. Then, the graph attention network (GAT) performs global weight calculations on these features to obtain accurate origami results. This method can not only effectively predict the two-dimensional structure of complex origami models such as paper cranes, but also provide detailed morphological predictions at various stages of the folding process, thereby helping designers optimize the origami process and improve design accuracy. Specifically, an origami prediction method based on multimodal graph learning includes the following implementation steps:
[0052] Step S1, obtaining folding design parameters and initial crease image;
[0053] Step S2, generating a graph topology structure and a two-dimensional crease image according to the folding design parameters;
[0054] Step S3, using a graph convolutional neural network to extract features from the graph topology and the two-dimensional fold image to obtain a first geometric feature; sending the initial fold image and the first geometric feature to a graph attention network for feature learning to obtain a second geometric feature and an image feature;
[0055] Step S4: Fusing the second geometric features and image features, analyzing the fused features using a multi-layer perceptron and outputting origami prediction results.
[0056] Based on the above process, the present invention can achieve efficient and high-precision prediction of origami results on the basis of improving the universality of the prediction model. In order to better understand the technical solution of the present invention, the specific implementation steps of the present invention are further explained and illustrated below.
[0057] Step S1, obtaining folding design parameters and initial crease image.
[0058] like Figure 3 As shown in the figure, the main input data used by the model include vertex coordinates, geometric constraints, crease edge sets, crease types, and folding degrees, etc. Among them, the geometric constraints are various geometric constraints in the folding theorem, and the crease types mainly include mountain creases, valley creases, and boundary creases.
[0059] The folding design parameters and the initial crease image are input by the user. Specifically, in this embodiment, vertex coordinates, geometric constraints, crease edge sets and crease types are selected as folding design parameters; among which, the constraints selected to be calculated are Maekawa theorem and Kawasaki theorem.
[0060] Furthermore, the Maekawa theorem states that in the folding direction of each fold that intersects at any vertex in the folding design diagram, the difference between the number of mountain folds and valley folds is always 2. The Maekawa theorem is expressed as:
[0061] ;
[0062] in, represents the set of all vertices in the crease design diagram, Represents a vertex i The number of creases at the mountain, Represents a vertex i The number of valley creases.
[0063] Furthermore, Kawasaki's theorem states that when the angles at the folds are numbered around any vertex in the fold design diagram, the sum of the angles of the folds with odd numbers is equal to the sum of the angles of the folds with even numbers, and is always 180°. Kawasaki's theorem is expressed as:
[0064] ;
[0065] in, k Represents a vertex i The number of each fold, Indicates k The fold angle at the crease.
[0066] Step S2: Generate a graph topology structure and a two-dimensional crease image according to the folding design parameters.
[0067] Based on Python algorithm, the graph topology is generated and the two-dimensional crease image is drawn according to the vertex coordinates, crease edge set, geometric constraints and crease types of the two-dimensional crease graph; the vertex coordinates are used as node features, the geometric constraints and crease types are used as edge features and stored in the graph topology structure, and the crease edge set is used to distinguish the types of edge features. This helps to improve the model's ability to extract features and obtain more accurate prediction results.
[0068] Step S3, using a graph convolutional neural network to extract features from the graph topology and the two-dimensional fold image to obtain a first geometric feature. The initial fold image and the first geometric feature are sent to the graph attention network for feature learning to obtain a second geometric feature and an image feature, that is, the graph attention network performs global weight calculation and feature learning on the node features and edge features in the initial fold image and the first geometric feature to obtain a second geometric feature and an image feature.
[0069] Step S3-1: Before using the graph convolutional neural network and the graph attention network for origami prediction, the graph convolutional neural network and the graph attention network need to be trained.
[0070] Step S3-1-1: To ensure the accuracy of model training, augment the data used for training to increase the amount of data used for training.
[0071] Step S3-1-2, define the fusion loss function according to the prediction task, compare the prediction result of the fused feature with the true label and calculate the loss value; then, back-propagate the loss value to the graph convolutional neural network and the graph attention network based on the back-propagation algorithm; iteratively update the parameters in the graph convolutional neural network and the graph attention network based on the gradient descent algorithm.
[0072] Step S3-2: Using the trained graph convolutional neural network and graph attention network to perform origami prediction can be achieved through the following process:
[0073] Step S3-2-1, using a graph convolutional neural network to extract features from the graph topology and the two-dimensional fold image, including: setting the convolution step size, using the graph convolutional neural network to generate a subgraph starting from the current node of the graph topology with the step size as the distance, transferring information between the node features and edge features of the subgraph, and aggregating the feature information embodied by the two-dimensional fold image to obtain the first geometric feature, that is, setting the step size of the convolution layer in the graph convolutional neural network to , starting from the current node, search step length , extract the node features and edge features in the subgraph, and Figure 4 As shown, information is transferred between node features and edge features in the subgraph.
[0074] Step S3-2-2: Send the initial fold image and the first geometric features (node features and edge features after information aggregation) to the self-attention layer in the graph attention network. Figure 5 As shown, the Graph Attention Network (GAT) has a i, we can find the node with the highest correlation score globally and fuse it with the features of the node to obtain the second geometric features and image features.
[0075] During computation, GAT uses a shared learnable parameter matrix The nodes are linearly transformed to obtain , and then calculate the attention coefficient of the current node to its neighboring nodes, that is:
[0076] ;
[0077] in, represents the attention coefficient, represents the activation function, represents the learnable parameter matrix,
[0078] and Respectively represent nodes i and nodes j In order to make the attention coefficient In order to make different nodes comparable, the Softmax function needs to be used to normalize them to obtain the normalized attention coefficient. The specific calculation formula is:
[0079] ;
[0080] in, represents the normalized attention coefficient, Represents a node i The set of neighbor nodes of k represents all neighbor nodes and includes nodes in the self-attention mechanism i itself.
[0081] Step S4: Fusing the second geometric features and the image features, using a multi-layer perceptron to analyze the fused features and outputting origami prediction results.
[0082] Step S4-1: Perform feature aggregation based on the attention coefficient calculated in step 3-2, and the node i New feature representation It is obtained by weighted summing of the features of its neighbor nodes, namely:
[0083] ;
[0084] in, For Node i New feature representation.
[0085] Furthermore, the pooling layer as part of the graph attention network adopts self-attention pooling (Self-Attention Graph Pooling). This layer first passes through the convolutional layer and activation function in the graph convolutional neural network to obtain the self-attention score (Self-Attention Score) of each node; then, according to the sorting of the self-attention scores, the first n nodes with the highest scores are retained to form a mask; finally, the features of the specified nodes and the topological information between these nodes are retained according to the mask.
[0086] Step S4-2, perform feature fusion on the second geometric feature and the image feature, assuming that the image feature vector corresponding to the image feature is I, and the geometric feature vector corresponding to the second geometric feature is R; at the same time, and are the dimensions of the image feature and the second geometric feature, respectively.
[0087] The feature transformation uses the fully connected layer mapping to transform the dimension, so that the second geometric feature and the image feature have a more appropriate dimension and more accurate representation. The feature calculation formula after mapping is as follows:
[0088] ;
[0089] in, Represents the activation function, the activation function is . Convolution and activation are performed on the image features and respectively, and the weight matrix obtained by learning and the bias term , for the original image features I Perform linear transformation and apply activation function to obtain the transformed image features ; The weight matrix obtained through learning and the bias term , the topological structure characteristics of the graph Perform linear transformation and apply activation function to obtain the transformed image features .
[0090] Step S4-3: Use a multi-layer perceptron to analyze the fused features and output the origami prediction results.
[0091] Step S4-3-1, introduce an attention matrix to dynamically determine the importance weights of image features and second geometric features in the fusion process. Specifically, let the attention module be A, and the attention module A uses the transformed image features and the second geometric feature As input, output two attention scores and, namely:
[0092] ;
[0093] in, represents the attention map of the image, Attention to the topological structure of the graph; represents the attention operation, Represents image features, Represents the second geometric feature.
[0094] Furthermore, the attention module A is a simple multi-layer perceptron (MLP), which includes two input layers (for receiving image features and second geometric features respectively), one hidden layer and two output layers (for outputting attention scores of image features and second geometric features respectively). The hidden layer uses the tanh function for nonlinear transformation, and the output layer uses the softmax function for normalization.
[0095] Step S4-3-2: According to the calculated attention score, the image feature and the second geometric feature are weightedly fused to obtain the fused feature F ,Right now:
[0096] ;
[0097] In order to further enhance the feature fusion effect and promote the information interaction between the image features and the second geometric features, this module introduces an additional interaction mechanism in the fusion layer. Specifically, the element-wise product (Hadamard product) between the image features and the second geometric features is calculated and then the fully connected layer is transformed, that is, the feature after the element-wise product is set to , the features after transformation by the fully connected layer are ;in, represents the weight, represents the bias term. Then, The fusion feature obtained by weighted fusion F Add them together to get the processed fusion features ,have This multimodal interactive fusion method allows the model to learn more complex relationships between image features and second geometric features (geometric information features), which helps to improve the quality of fused features.
[0098] Step S4-3-3: The fused features are finally transformed through the output layer to meet the needs of subsequent prediction tasks. The output layer can be a simple fully connected layer that maps the fused features to the target dimension, as shown in the following example. Figure 6 As shown, let the output layer weight matrix be , the bias vector is , the final output fusion feature vector for:
[0099] ;
[0100] Use multi-layer perceptron to fused features (fused feature vector ) to analyze and output origami prediction results; Figure 7 The example picture shown in the figure, after the fold prediction, the output prediction result is as follows Figure 8 Based on the above process, the present invention can achieve efficient and high-precision prediction of origami results and improve the universality of the prediction model.
[0101] Embodiment 2
[0102] This embodiment discloses an origami prediction system based on multimodal graph learning.
[0103] A paper-folding prediction system based on multimodal graph learning, comprising:
[0104] The data acquisition module is configured to: acquire folding design parameters and initial crease images;
[0105] An object generation module is configured to: generate a graph topology structure and a two-dimensional crease image according to the folding design parameters;
[0106] The feature extraction module is configured to: use a graph convolutional neural network to extract features from the graph topology and the two-dimensional fold image to obtain a first geometric feature; send the initial fold image and the first geometric feature to a graph attention network for feature learning to obtain a second geometric feature and an image feature;
[0107] The origami prediction module is configured to: perform feature fusion on the second geometric feature and the image feature, analyze the fused features using a multi-layer perceptron and output an origami prediction result.
[0108] Embodiment 3
[0109] The purpose of this embodiment is to provide a computer-readable storage medium.
[0110] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in an origami prediction method based on multimodal graph learning as described in Embodiment 1 of the present disclosure.
[0111] Embodiment 4
[0112] The purpose of this embodiment is to provide an electronic device.
[0113] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of an origami prediction method based on multimodal graph learning as described in the first embodiment of the present disclosure are implemented.
[0114] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0115] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0116] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A paper-folding prediction method based on multimodal graph learning, characterized in that: include: Obtain folding design parameters and initial crease images; generating a graph topology structure and a two-dimensional crease image according to the folding design parameters; Using a graph convolutional neural network to extract features from the graph topology and the two-dimensional crease image to obtain a first geometric feature; Sending the initial crease image and the first geometric feature to a graph attention network for feature learning to obtain a second geometric feature and an image feature; The second geometric feature and the image feature are fused, and the fused features are analyzed using a multi-layer perceptron to output an origami prediction result.
2. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that: The folding design parameters include vertex coordinates, geometric constraints, crease edge sets and crease types.
3. The origami prediction method based on multimodal graph learning according to claim 2, characterized in that: Vertex coordinates are used as node features, and geometric constraints and crease types are used as edge features.
4. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that: A graph convolutional neural network is used to extract features from a graph topology structure and a two-dimensional fold image, including: setting a convolution step size, using a graph convolutional neural network to generate a subgraph starting from a current node of the graph topology structure with a step size as a distance, transmitting information between the node features and edge features of the subgraph, and aggregating the feature information embodied by the two-dimensional fold image to obtain a first geometric feature.
5. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that: The graph attention network performs global weight calculation and feature learning on the node features and edge features in the initial crease image and the first geometric features to obtain the second geometric features and image features.
6. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that: Before using graph convolutional neural networks and graph attention networks for origami prediction, they need to be trained; to ensure the accuracy of model training, the data used for training is augmented to increase the amount of data used for training.
7. The origami prediction method based on multimodal graph learning according to claim 6, characterized in that: The graph convolutional neural network and the graph attention network are trained, including: defining a fusion loss function according to a prediction task, comparing the prediction results of the fused features with the true labels and calculating the loss values; then, back-propagating the loss values to the graph convolutional neural network and the graph attention network based on a back-propagation algorithm; and iteratively updating the parameters in the graph convolutional neural network and the graph attention network based on a gradient descent algorithm.
8. A paper-folding prediction system based on multimodal graph learning, characterized in that: include: The data acquisition module is configured to: acquire folding design parameters and initial crease images; An object generation module is configured to: generate a graph topology structure and a two-dimensional crease image according to the folding design parameters; A feature extraction module is configured to: use a graph convolutional neural network to extract features from the graph topology structure and the two-dimensional fold image to obtain a first geometric feature; Sending the initial crease image and the first geometric feature to a graph attention network for feature learning to obtain a second geometric feature and an image feature; The origami prediction module is configured to: perform feature fusion on the second geometric feature and the image feature, analyze the fused features using a multi-layer perceptron and output an origami prediction result.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of an origami prediction method based on multimodal graph learning as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the origami prediction method based on multimodal graph learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Textile crease detection method based on deep learning
CN111583253A
Silencer design method based on paper folding structure and topological optimization
CN118747461A