Origami prediction method and system based on multimodal graph learning

JP2026142494AActive Publication Date: 2026-09-07SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025105732
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2025-06-23
Publication Date
2026-09-07
Estimated Expiration
2045-06-23

AI Technical Summary

Benefits of technology

【0022】 以上の1つ又は複数の技術的手段は以下の有益な効果がある。 (1)本発明は、グラフ畳み込みニューラルネットワークを用いてグラフトポロジー構造と二次元折り線画像に対して特徴抽出を行うことで、幾何的特徴を取得し、また、グラフ注意ネットワークを用いて幾何的特徴に対して更なる特徴学習を行う時に、画像特徴を学習でき、最後に学習した画像特徴と幾何的特徴を融合する。従って、本発明は、マルチモーダル特徴を正確に融合しながら、折り紙定理中の制約条件を満たすことができ、異なるタイプの折り紙シーンでそれぞれ良好な適応性を示し、このネットワークアーキテクチャの設計によって、各種の折り紙シーンの挑戦に対応することが可能になり、予測モデルの汎用性が向上する。 (2)本発明は、グラフ畳み込みニューラルネットワークとグラフ注意ネットワークを組み合わせることで、キー情報を保ちながら、計算量を有効に減らすことができ、同時に、モデルは優れた汎化能力を有し、類似構造の場合にコンピュータの計算力消費を大幅に低減させ、計算効率を向上させる。 (3)本発明は幾何的構造特徴を抽出するだけでなく、画像特徴からも分析することで、形状についても、具体的な幾何的構造についても予測結果が高い精度を有することを確保できる。同時に、本発明はトレーニング用のデータに対して拡張操作を行い、トレーニング用のデータ量を増加させることで、トレーニングしたモデルにより好適な予測能力を持たせる。従って、本発明は、折り紙結果の高精度な予測を実現することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026142494000001_ABST
    Figure 2026142494000001_ABST
Patent Text Reader

Abstract

This invention provides an origami prediction method and system based on multimodal graph learning. [Solution] In this invention, folding design parameters and an initial fold line image are acquired, a graph topology structure and a two-dimensional fold line image are generated using the folding design parameters, feature extraction is performed on the graph topology structure and the two-dimensional fold line image using a graph convolutional neural network to acquire first geometric features, the initial fold line image and the first geometric features are transmitted to a graph attention network for feature learning, second geometric features and image features are acquired, feature fusion is performed on the second geometric features and image features, and the fused features are analyzed using a multilayer perceptron to output an origami prediction result. As a result, this invention can achieve efficient and highly accurate prediction of origami results and improve the versatility of the prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] (Cross-Reference to Related Applications) The present invention claims priority from the Chinese patent application filed with the China National Intellectual Property Administration on February 26, 2025, with application number 202510213165.5 and the title of the invention being "Origami prediction method and system based on multimodal graph learning", the entire content of which is incorporated into the present invention by reference and constitutes a part of the present invention for all purposes.

[0002] The present invention belongs to the technical field of image processing, and particularly relates to an origami prediction method and system based on multimodal graph learning. [[Background Art]]

[0003] The description in this section merely provides background technical information related to the present invention, and does not necessarily constitute prior art.

[0004] An origami structure is a novel structural system that does not require cutting or stretching, and can be folded along preset creases (mountain creases, valley creases, etc.) based on a two-dimensional plane. The origami structure has the characteristics of changeable modality, high designability, excellent rigidity and folding expansion ratio, and exhibits good performance in device structures. Therefore, origami structures are widely used in aerospace, medical devices, robotic technology and other fields. For example, a satellite solar panel only has a diameter of 2.5m when folded, and can reach a diameter of 27m when fully deployed. It has a large folding expansion ratio, and can obtain a larger practical area with a small storage space.

[0005] The key to origami prediction lies in simulating and predicting the folding process and final result based on initial conditions of origami, such as the size, shape, and fold line distribution of the materials to be folded, while being constrained by folding rules. By calculating and inferring the three-dimensional folded shape from two-dimensional fold line images, it is possible not only to predict the results of artistic origami but also to leverage the advantages of origami structures in industrial fields.

[0006] The inventor discovered that conventional methods for predicting results based on origami structures have several technical problems. For example, the following problems exist: (1) Current common methods for achieving origami results rely primarily on manual experience and finite simulation reasoning. For example, a designer first designs the fold lines and then implements them manually or mechanically according to the predetermined fold lines. However, such methods are difficult to implement on a large scale because they involve long working periods, large workloads, and require extremely high precision in manufacturing equipment when designing relatively precise foldable structures. (2) Conventional origami design methods primarily rely on traditional origami rules such as the Huzita-Hatori axiom, the Maekawa condition, Maekawa's Theorem, and the 2-Colorable condition. While these theories clearly constrain the geometric transformations of the folding process, the computational processes involved are extremely complex. Consequently, computational efficiency is low. (3) Conventional simulation prediction software (e.g., Origami Simulator) that can perform inference calculations for the folding process of origami structures does accelerate the prediction speed to some extent. However, such simulation software needs to be recalculated every time a prediction is made, lacks learning generalization ability, and is severely limited in the case of complex origami structures. At the same time, there is a growing amount of research seeking new origami prediction methods using deep learning techniques, but current deep learning-based models have difficulty simultaneously capturing geometric and shape features in origami structures when used in the field of origami result prediction. As a result, the prediction results obtained are not accurate in terms of both modality and geometric structure, and the prediction results are undesirable.

[0007] In summary, existing origami prediction methods generally suffer from problems such as low versatility, low prediction efficiency, and low prediction accuracy when dealing with large-scale and highly complex prediction tasks. [Overview of the Initiative]

[0008] To overcome the shortcomings of the conventional technology described above, the present invention provides an origami prediction method and system based on multimodal graph learning that can achieve efficient and highly accurate prediction of origami results and improve the versatility of the prediction model.

[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical means.

[0010] In a first embodiment of the present invention, a method for predicting origami based on multimodal graph learning is provided.

[0011] Steps include obtaining folding design parameters and initial fold line images, The steps include generating a graph topology structure and a two-dimensional fold line image using the aforementioned folding design parameters, The process involves using a graph convolutional neural network to extract features from the graph topology structure and the two-dimensional folded line image, obtaining first geometric features, transmitting the initial folded line image and the first geometric features to a graph attention network for feature learning, and obtaining second geometric features and image features. This is an origami prediction method based on multimodal graph learning, which includes the steps of performing feature fusion on the second geometric feature and the image feature, analyzing the fused features using a multilayer perceptron, and outputting an origami prediction result.

[0012] Furthermore, the folding design parameters include vertex coordinates, geometric constraints, a set of folded edge lines, and a type of folded line.

[0013] Furthermore, vertex coordinates are used as node features, and geometric constraints and fold line types are used as edge features.

[0014] Furthermore, performing feature extraction on graph topology structures and two-dimensional foldline images using a graph convolutional neural network involves setting the convolution stride, generating subgraphs from the current node of the graph topology structure using the stride as the distance, transferring information to the node and edge features of the subgraphs, and aggregating the feature information represented by the two-dimensional foldline image to obtain the first geometric feature.

[0015] Furthermore, the graph attention network performs global weight calculation and feature learning on the initial polyline image and the node and edge features of the first geometric feature to obtain the second geometric feature and image features.

[0016] Furthermore, before performing origami prediction using graph convolutional neural networks and graph attention networks, it is necessary to train these networks. To ensure the accuracy of model training, augmentation operations are performed on the training data to increase the amount of training data.

[0017] Furthermore, training a graph convolutional neural network and a graph attention network involves defining a fusion loss function by a prediction task, calculating a loss value by comparing the predicted results of the fused features with the correct labels, then backpropagating the loss value to the graph convolutional neural network and the graph attention network based on a backpropagation algorithm, and repeatedly updating the parameters in the graph convolutional neural network and the graph attention network based on a gradient descent algorithm.

[0018] In a second aspect of the present invention, an origami prediction system based on multimodal graph learning is provided.

[0019] A data acquisition module configured to acquire folding design parameters and an initial fold line image, A target generation module configured to generate a graph topology structure and a two-dimensional folded line image using the aforementioned folding design parameters, A feature extraction module is configured to perform feature extraction on the graph topology structure and two-dimensional fold line image using a graph convolutional neural network, obtain a first geometric feature, transmit the initial fold line image and the first geometric feature to a graph attention network for feature learning, and obtain a second geometric feature and image features. This is an origami prediction system based on multimodal graph learning, comprising: an origami prediction module configured to perform feature fusion on the second geometric feature and image feature, analyze the fused features using a multilayer perceptron, and output an origami prediction result.

[0020] In a third aspect of the present invention, a computer-readable storage medium is provided, which stores a program, and when the program is executed by a processor, the steps in the origami prediction method based on multimodal graph learning described in the first aspect of the present invention are realized.

[0021] A fourth aspect of the present invention provides an electronic device comprising a memory, a processor, and a program stored in the memory and executable by the processor, wherein when the program is executed by the processor, the steps in the origami prediction method based on multimodal graph learning described in the first aspect of the present invention are realized.

[0022] One or more of the above technical measures have the following beneficial effects: (1) The present invention uses a graph convolutional neural network to extract features from graph topology structures and two-dimensional folded line images to acquire geometric features, and also uses a graph attention network to learn image features when further feature learning is performed on the geometric features, and finally the learned image features and geometric features are fused together. Therefore, the present invention can satisfy the constraints in the origami theorem while accurately fusing multimodal features, and shows good adaptability to different types of origami scenes. The design of this network architecture makes it possible to respond to the challenges of various origami scenes and improves the versatility of the predictive model. (2) The present invention combines a graph convolutional neural network and a graph attention network to effectively reduce computational complexity while preserving key information. At the same time, the model has excellent generalization capabilities, significantly reducing the computational power consumption of the computer in the case of similar structures and improving computational efficiency. (3) The present invention not only extracts geometric structural features but also analyzes them from image features, thereby ensuring that the prediction results have high accuracy for both shape and specific geometric structure. At the same time, the present invention performs an augmentation operation on the training data, increasing the amount of training data, thereby giving the trained model more suitable predictive capabilities. Accordingly, the present invention can achieve highly accurate prediction of origami results.

[0023] The advantages of additional aspects of the present invention are, in part, set forth in the following description, in part, become apparent from the following description, or are understood through the practice of the present invention. [Brief explanation of the drawing]

[0024] The specification and drawings forming part of the present invention are intended for further understanding of the present invention, and the exemplary embodiments of the present invention and the descriptions thereof are for the purpose of interpreting the present invention, and are not intended to inappropriately limit the present invention.

[0025] [Figure 1] It is a flowchart of an origami prediction method based on multimodal graph learning in Embodiment 1 of the present invention. [Figure 2] It is a schematic diagram of data relationships of the origami prediction method based on multimodal graph learning in Embodiment 1 of the present invention. [Figure 3] It is a schematic diagram of model input data in Embodiment 1 of the present invention. [Figure 4] It is a schematic diagram of the information transmission process of a graph convolutional neural network in Embodiment 1 of the present invention. [Figure 5] It is a schematic diagram of the information transmission process of a graph attention network in Embodiment 1 of the present invention. [Figure 6] It is a schematic diagram of a feature fusion process in Embodiment 1 of the present invention. [Figure 7] It is an exemplary diagram of a two-dimensional crease pattern design drawing in Embodiment 1 of the present invention. [Figure 8] It is a schematic diagram of crease result prediction corresponding to a two-dimensional crease pattern design drawing in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0026] It should be noted that the following detailed description is exemplary and is intended to further illustrate the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as generally understood by those skilled in the art to which the present invention pertains.

[0027] It should be noted that the technical terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.

[0028] The embodiments and features of the present invention can be combined with each other, as long as they do not contradict each other. Example 1

[0029] This embodiment discloses an origami prediction method based on multimodal graph learning.

[0030] As shown in Figures 1 and 2, this is an origami prediction method based on multimodal graph learning that can be applied to predicting the two-dimensional folding structure of complex origami models, such as the folding pattern of an origami crane. By inputting folding design parameters such as the vertex coordinates, fold line types, and geometric constraints of the origami crane, the model generates a corresponding graph topology structure, and a graph convolutional neural network (GCN) can extract features from the fold line image and geometric information. Next, a graph attention network (GAT) performs global weight calculation on these features to obtain accurate origami results. This method can not only effectively predict the two-dimensional structure of complex origami models such as origami cranes, but can also provide detailed morphological predictions at each stage of the folding process, thereby helping designers optimize the origami process and improve design accuracy. Specifically, Step S1 involves obtaining folding design parameters and an initial fold line image, Step S2 generates a graph topology structure and a two-dimensional fold line image using the aforementioned folding design parameters, Step S3 involves using a graph convolutional neural network to extract features from the graph topology structure and the two-dimensional folded line image, obtaining a first geometric feature, transmitting the initial folded line image and the first geometric feature to a graph attention network for feature learning, and obtaining a second geometric feature and image features. This is an origami prediction method based on multimodal graph learning, which includes step S4, which involves performing feature fusion on the second geometric feature and the image feature, and then analyzing the fused features using a multilayer perceptron to output an origami prediction result.

[0031] According to the above process, the present invention can improve the versatility of the prediction model and achieve efficient and highly accurate prediction of origami results. To better understand the technical means of the present invention, the specific steps for implementing the present invention will be further interpreted and explained below.

[0032] In step S1, the folding design parameters and the initial fold line image are obtained.

[0033] As shown in Figure 3, the main input data used in the model includes vertex coordinates, geometric constraints, a set of folded edge lines, folded line types, and the degree of folding. Here, the geometric constraints are the various geometric constraints in the folding theorem, and the folded line types mainly include mountain folds, valley folds, and boundary folds.

[0034] The user inputs the folding design parameters and initial fold line image themselves. Specifically, in this embodiment, the user selects vertex coordinates, geometric constraints, a set of fold line edges, and a fold line type as folding design parameters. The computationally necessary constraints selected are the Maekawa theorem and the Kawasaki theorem.

[0035] Furthermore, the Maekawa theorem states that in a fold line design, the difference between the number of mountain folds and valley folds in the folding direction of each fold line intersecting any vertex is always 2. The Maekawa theorem is,

number

number

number

number

[0036] Furthermore, the Kawasaki theorem states that if the included angles of the fold lines around any vertex in a fold line design are numbered, then the sum of the included angles of the odd-numbered fold lines is equal to the sum of the included angles of the even-numbered fold lines, and is always 180°.

number

number

[0037] In step S2, a graph topology structure and a two-dimensional fold line image are generated using folding design parameters.

[0038] Based on a Python algorithm, a graph topology structure is generated from the vertex coordinates, fold line edge set, geometric constraints, and fold line type of a two-dimensional fold line image. The two-dimensional fold line image is then plotted. The vertex coordinates are stored as node features, and the geometric constraints and fold line type are stored as edge features in the graph topology structure. The fold line edge set is used to distinguish between different types of edge features. This contributes to improving the model's feature extraction capabilities and obtaining more accurate prediction results.

[0039] In step S3, a graph convolutional neural network is used to extract features from the graph topology structure and the two-dimensional fold line image to obtain the first geometric features. The initial fold line image and the first geometric features are sent to the graph attention network for feature learning to obtain the second geometric features and image features. Specifically, the graph attention network performs global weight calculation and feature learning on the node and edge features in the initial fold line image and the first geometric features to obtain the second geometric features and image features.

[0040] In step S3-1, before performing origami prediction using the graph convolutional neural network and graph attention network, it is necessary to train the graph convolutional neural network and graph attention network.

[0041] In step S3-1-1, to ensure the accuracy of model training, an augmentation operation is performed on the training data to increase the amount of training data.

[0042] In step S3-1-2, a fusion loss function is defined by the prediction task, the predicted result of the fused features is compared with the ground truth label to calculate the loss value, and then the loss value is backpropagated to the graph convolutional neural network and the graph attention network based on the backpropagation algorithm, and the parameters in the graph convolutional neural network and the graph attention network are repeatedly updated based on the gradient descent algorithm.

[0043] In step S3-2, origami prediction is performed using a trained graph convolutional neural network and a graph attention network, which can be achieved through the following steps.

[0044] Step S3-2-1 involves performing feature extraction on the graph topology structure and the two-dimensional fold line image using a graph convolutional neural network, which includes setting the convolution stride, generating a subgraph from the current node of the graph topology structure using the stride as the distance, transferring information to the node and edge features of the subgraph, and aggregating the feature information represented by the two-dimensional fold line image to obtain the first geometric feature, that is, setting the stride of the convolutional layer in the graph convolutional neural network

number

number

[0045] In step S3-2-2, the initial pultex image and the first geometric features (node ​​features and edge features after information aggregation) are transmitted to the Self-Attention Layer of the graph attention network. As shown in Figure 5, the graph attention network (GAT) globally finds the node with the highest correlation score for each node i, fuses it with the features of that node, and obtains the second geometric features and image features.

[0046] When performing calculations, GAT uses a shared, learnable parameter matrix.

number

number

number

number

number

number

number

number

number

number

number

number

[0047] In step S4, feature fusion is performed on the second geometric feature and the image feature, and the fused features are analyzed using a multilayer perceptron to output the origami prediction result.

[0048] In step S4-1, feature aggregation is performed using the attention coefficient calculated in step 3-2, and a new feature representation of node i is created.

number

number

number

[0049] Furthermore, the pooling layer becomes part of the graph attention network, employing Self-Attention Graph Pooling. This layer first processes the data using convolutional layers and activation functions in a graph convolutional neural network to obtain a Self-Attention Score for each node. Next, it sorts the Self-Attention Scores to retain the n nodes with the highest scores from the previous step, forming a mask. Finally, the mask retains the features of the specified nodes and the topological information between these nodes.

[0050] In step S4-2, feature fusion is performed on the second geometric feature and the image feature, with the image feature vector corresponding to the image feature being denoted as I, and the geometric feature vector corresponding to the second geometric feature being denoted as R, and simultaneously,

number

number

[0051] In feature transformation, dimensional transformation is performed by fully connected layer mapping, giving the second geometric features and image features appropriate dimensions and more accurate representativeness. The formula for calculating the features after mapping is:

number

number

number

number

number

number

number

number

number

number

[0052] In step S4-3, a multilayer perceptron is used to analyze the features after fusion and output the origami prediction result.

[0053] In step S4-3-1, an attention matrix is ​​introduced to dynamically determine the importance weights in the fusion process of image features and second geometric features. Specifically, the attention module is A, and attention module A is the transformed image feature.

number

number

number

number

number

number

number

number

[0054] Furthermore, attention module A is a simple multilayer perceptron (MLP) and includes two input layers (one for receiving image features and the other for receiving second geometric features), one hidden layer, and two output layers (one for outputting attention scores for image features and the other for receiving second geometric features). The hidden layer performs a nonlinear transformation using the tanh function, and the output layer performs normalization using the softmax function.

[0055] In step S4-3-2, the calculated attention score is used to perform weighted fusion on the image features and the second geometric features to obtain the fused feature F. That is, the calculation formula is:

number

number

number

number

number

number

number

number

[0056] In step S4-3-3, the fused features are ultimately transformed through an output layer to adapt to the requirements of the subsequent prediction task. The output layer may be a simple fully connected layer that maps the fused features to the target dimension, specifically as shown in Figure 6, where the weight matrix of the output layer is

number

number

number

number

number

[0057] This embodiment discloses an origami prediction system based on multimodal graph learning.

[0058] A data acquisition module configured to acquire folding design parameters and an initial fold line image, A target generation module configured to generate a graph topology structure and a two-dimensional folded line image using the aforementioned folding design parameters, A feature extraction module is configured to perform feature extraction on the graph topology structure and two-dimensional fold line image using a graph convolutional neural network, obtain a first geometric feature, transmit the initial fold line image and the first geometric feature to a graph attention network for feature learning, and obtain a second geometric feature and image features. This is an origami prediction system based on multimodal graph learning, comprising: an origami prediction module configured to perform feature fusion on the second geometric feature and image feature, analyze the fused features using a multilayer perceptron, and output an origami prediction result. Example 3

[0059] The objective of this embodiment is to provide a computer-readable storage medium.

[0060] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, the steps in the multimodal graph learning-based origami prediction method described in Embodiment 1 of this disclosure are realized. Example 4

[0061] The objective of this embodiment is to provide an electronic device.

[0062] The electronic device comprises memory, a processor, and a program stored in the memory and executable by the processor, wherein when the program is executed by the processor, the steps in the multimodal graph learning-based origami prediction method described in Embodiment 1 of this disclosure are realized.

[0063] Each step involving the apparatus in Examples 2, 3, and 4 corresponds to Example 1 of the Method, and for specific embodiments, refer to the relevant descriptive sections of Example 1. The term “computer-readable storage medium” should be understood as a single or multiple medium containing one or more command sets, and further understood as including any medium, any medium which stores, encodes, or carries command sets to be executed by a processor, and causes the processor to execute any method of the present invention.

[0064] As those skilled in the art will see, each module or step of the present invention described above can be implemented on a general-purpose computer device, and selectively, they can be implemented as program code executable by the computer device, so they can be stored in a memory device and executed by the computer device, or they can be manufactured as individual integrated circuit modules, or multiple modules or steps can be manufactured as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0065] Although specific embodiments of the present invention have been described above with reference to the drawings, this does not limit the scope of protection of the present invention. As those skilled in the art will understand, various modifications and variations made by those skilled in the art based on the technical means of the present invention without requiring any creative effort are also included within the scope of protection of the present invention.

Claims

1. Steps include obtaining folding design parameters and initial fold line images, The steps include generating a graph topology structure and a two-dimensional fold line image using the aforementioned folding design parameters, The process involves using a graph convolutional neural network to extract features from the graph topology structure and the two-dimensional folded line image, obtaining first geometric features, transmitting the initial folded line image and the first geometric features to a graph attention network for feature learning, and obtaining second geometric features and image features. A computer-based multimodal graph learning method for predicting origami, comprising the steps of: performing feature fusion on the second geometric feature and the image feature, analyzing the fused features using a multilayer perceptron, and outputting an origami prediction result.

2. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that the folding design parameters include vertex coordinates, geometric constraints, a set of fold line edges, and a type of fold line.

3. The origami prediction method based on multimodal graph learning according to claim 2, characterized in that vertex coordinates are used as node features and geometric constraints and fold line types are used as edge features.

4. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that the feature extraction of a graph topology structure and a two-dimensional folded line image using a graph convolutional neural network includes setting a convolution stride, generating a subgraph from the current node of the graph topology structure using the stride as the distance using the graph convolutional neural network, transferring information to the node features and edge features of the subgraph, and aggregating the feature information represented by the two-dimensional folded line image to obtain a first geometric feature.

5. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that the graph attention network performs global weight calculation and feature learning on the initial fold line image and node and edge features in the first geometric feature to obtain the second geometric feature and image feature.

6. The origami prediction method based on multimodal graph learning according to claim 1, characterized in that, before performing origami prediction using a graph convolutional neural network and a graph attention network, the graph convolutional neural network and the graph attention network must be trained, and in order to ensure the accuracy of model training, an augmentation operation is performed on the training data to increase the amount of training data.

7. The origami prediction method based on multimodal graph learning according to claim 6, characterized in that training a graph convolutional neural network and a graph attention network includes defining a fusion loss function by a prediction task, calculating a loss value by comparing the predicted result of the fused features with the correct label, then backpropagating the loss value to the graph convolutional neural network and the graph attention network based on a backpropagation algorithm, and repeatedly updating the parameters in the graph convolutional neural network and the graph attention network based on a gradient descent algorithm.

8. A data acquisition module configured to acquire folding design parameters and an initial fold line image, A target generation module configured to generate a graph topology structure and a two-dimensional folded line image using the aforementioned folding design parameters, A feature extraction module is configured to perform feature extraction on the graph topology structure and two-dimensional fold line image using a graph convolutional neural network, obtain a first geometric feature, transmit the initial fold line image and the first geometric feature to a graph attention network for feature learning, and obtain a second geometric feature and image features. An origami prediction system based on multimodal graph learning, comprising: an origami prediction module configured to perform feature fusion on the second geometric feature and the image feature, and to analyze the fused features using a multilayer perceptron to output an origami prediction result.

9. A computer-readable storage medium in which a program is stored, wherein when the program is executed by a processor, the steps in the origami prediction method based on multimodal graph learning described in any one of claims 1 to 7 are realized.

10. An electronic device comprising memory, a processor, and a program stored in memory and executable by the processor, wherein when the program is executed by the processor, the steps in the origami prediction method based on multimodal graph learning described in any one of claims 1 to 7 are realized.