Method for generating component tree by assembling picture specification based on large language model

By identifying and separating non-component legends in the assembly picture manual and using a large language model to generate a step tree, the difficulty of understanding the assembly sequence in the assembly picture manual is solved, and precise modeling of complex assembly and automated support for machine assembly is achieved.

CN120014346APending Publication Date: 2025-05-16WUXI DYNAVISION MIYAHARA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510092579.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generate a component tree that assembles a picture manual, especially when facing complex image and text input, the application of multimodal large models has limitations.

Method used

By identifying and separating non-component legends in the assembly picture instruction manual, retaining component information, and using a large language model to organize the assembly picture into a step tree based on the few sample prompts, and generating a component tree based on the component information.

Benefits of technology

Effectively extract the component information used in each step in the manual and the interdependence between the steps, improving the accuracy of understanding the assembly sequence and supporting the automation of machine assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014346A_ABST
    Figure CN120014346A_ABST
Patent Text Reader

Abstract

The invention provides a method for generating a component tree by assembling a picture specification based on a large language model. The method for generating the component tree through the assembly picture specification based on the large language model comprises the steps that non-component legends in the assembly picture specification are recognized and separated, and component information is reserved; the large language model arranges the assembly pictures into a step tree according to the small sample prompt, and the step tree displays assembly steps through a tree-shaped structure; and combining the component information and the step tree to generate a component tree which displays the premise and subsequent relationship between the components through a tree structure. According to the method, the problem of difficulty in extracting components used for understanding the assembly sequence in the assembly picture specification is effectively solved, the component information used in each step in the specification and the mutual dependency relationship between the steps are effectively extracted, and the method has important significance on machine assembly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision technology and large language models, and in particular to a method for generating a component tree from an assembly picture instruction manual based on a large language model. Background Art

[0002] In recent years, the development of embodied intelligence has gradually been deeply integrated with other technologies. In particular, with the improvement of computing power and the advancement of sensor technology, the performance of robots in perception, navigation, and motion planning has been significantly improved. As an important task of embodied intelligence, the automated assembly of assemblies usually includes two stages: understanding and execution. Among them, the breakthrough in task understanding ability is due to the rapid development of computer vision and natural language processing technology. Since the breakthrough of deep learning in the ImageNet Challenge in 2012, these technical fields have made great progress, laying the foundation for the complex understanding of assembly tasks.

[0003] Deep learning architectures represented by Transformer have further promoted the leapfrog development of computer vision and natural language processing. Transformer has greatly improved the model's ability to handle long-distance dependencies through the self-attention mechanism. Large language models represented by the GPT series have performed well in tasks such as text generation, question answering, and language translation. Its subsequent development has further expanded the robot's ability to process multimodal information. In particular, models such as GPT-4 enable robots to make efficient decisions when faced with complex tasks by integrating multimodal information such as text and images. The development of these technologies provides strong technical support for the understanding of key contextual relationships and the connection of steps in assembly tasks.

[0004] In modern manufacturing, complex product assembly involves the automated assembly process of many parts; in home environments, scenarios such as furniture assembly place higher demands on robots. Among these tasks, generating assembly step trees and component trees is one of the core links in achieving automated assembly. The component tree not only clarifies the hierarchical relationship between the parts, but also provides a basis for the construction and optimization of the assembly process. In recent years, technological breakthroughs in deep learning and multimodal large models have enabled intelligent agents to understand the contextual logic of assembly tasks more comprehensively, which provides new possibilities for component tree generation. However, in the face of complex image and text inputs, the practical application of multimodal large models still has certain limitations. In order to further improve the accuracy and applicability of the model, it may be necessary to combine optimization methods such as input preprocessing and few-sample reference to better meet the needs in real scenarios.

[0005] Therefore, it is necessary to propose a new method for generating a component tree from assembly picture instructions based on a large language model. Summary of the invention

[0006] One of the purposes of the present invention is to overcome at least one deficiency in the prior art and to provide a method for generating a component tree from an assembly picture description based on a large language model.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0008] According to a first aspect of the present invention, a method for generating a component tree from an assembly picture instruction based on a large language model is provided. The method for generating a component tree from an assembly picture instruction based on a large language model comprises:

[0009] Identify and separate non-component legends in assembly picture instructions and retain component information;

[0010] The large language model organizes the assembly picture into a step tree based on the few sample prompts, wherein the step tree displays the assembly steps in a tree structure;

[0011] The component information and the step tree are combined to generate a component tree, wherein the component tree displays the prerequisite and successor relationships between components through a tree structure.

[0012] Optionally, the identifying and separating non-component legends in the assembly picture instructions and retaining component information includes:

[0013] Regular legends with regular geometric shapes as boundaries are identified through computer vision methods.

[0014] Optionally, after identifying the regular legend with a regular geometric shape as a boundary by a computer vision method, the method further includes:

[0015] The recognized rule legend is compared with the mask generated by the segmentation method, and the mask whose overlap rate reaches the set value is determined as the final legend recognition result.

[0016] Optionally, the identifying a regular legend with a regular geometric shape as a boundary by a computer vision method includes:

[0017] Identify a circular frame with a circle as the boundary by using the Hough circle detection method; and / or,

[0018] The Hough line detection method is used to identify straight lines and screen their angles. The distance between horizontal and vertical straight line breakpoints is calculated to determine whether the two intersect and infer whether they belong to the same rectangular frame. Based on this, a rectangular frame legend with a rectangle as the boundary is defined.

[0019] Optionally, parameter settings in the Hough circle detection method are set according to the size of the assembly picture.

[0020] Optionally, the identifying and separating non-component legends in the assembly picture instructions and retaining component information includes:

[0021] Using the object detection model to frame auxiliary schematic examples with irregular shapes but clear meaning; and / or,

[0022] The step number legend is identified by optical character recognition.

[0023] Optionally, the auxiliary schematic diagram example includes at least an arrow and / or a hand-shaped indicator image appearing in the assembly picture, which is used to indicate the assembler's operation.

[0024] Optionally, the specific coverage of the non-component legend is extracted and separated through an image segmentation model, retaining the component information.

[0025] Optionally, the large language model organizes the assembly picture into a step tree according to the few sample prompts, including:

[0026] Using the multimodal large model GPT-4o and the few-shot learning method, the input assembly images are sorted according to the given step tree examples to generate a step tree.

[0027] Optionally, the root node of the component tree represents the entire structure after assembly;

[0028] The component tree indicates that at least two prerequisite components are assembled in sequence to form subsequent components until the overall structure is assembled.

[0029] Different from the prior art, in the method for generating a component tree from an assembly picture instruction manual based on a large language model provided by the present invention, component information is retained by identifying and separating non-component legends in the assembly picture instruction manual. The large language model organizes the assembly picture into a step tree based on a few sample prompts, and the component tree can be generated by combining the component information and the step tree. This effectively solves the difficulty of extracting the components used for understanding the assembly sequence in the assembly picture instruction manual, and effectively extracts the component information used in each step in the instruction manual, as well as the interdependence between the steps, which is of great significance to machine assembly.

[0030] By using traditional computer vision methods, detection models, and optical character recognition methods to jointly detect various forms of legends, and using segmentation models to remove non-component legends and retain component information, the problem of legend diversity is solved and the component information in the picture description is highlighted.

[0031] The tree-like sequential relationship between the steps generated by the few-shot learning of the large language model is combined to finally obtain the component tree. By utilizing the prior knowledge and understanding ability of the large model, the problem of difficulty in understanding the sequential relationship of complex picture instructions is solved, the accuracy of the analysis is improved, and it is helpful for subsequent machine assembly.

[0032] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of a method for generating a component tree from an assembly picture description based on a large language model provided by an embodiment of the present invention.

[0034] Figure 2 A schematic flow chart of a method for generating a component tree provided in one embodiment of the present invention.

[0035] Figure 3 A schematic diagram of component information and non-component legends in an assembly picture instruction manual provided in an embodiment of the present invention.

[0036] Figure 4 A schematic diagram of arranging an assembly picture into a step tree provided in an embodiment of the present invention.

[0037] Figure 5 A schematic diagram of a component tree provided for an embodiment of the present invention.

[0038] in:

[0039] 12-Rectangular frame legend;

[0040] 16-circular frame legend;

[0041] 32- auxiliary schematic diagram example;

[0042] 36-Step number legend;

[0043] 50-parts;

[0044] 92-Assembly picture;

[0045] 94-step tree. DETAILED DESCRIPTION

[0046] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0048] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0049] It should be noted that the directional words such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described at the angles shown in the drawings and should not be understood as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is connected to another element "upper" or "lower", it can not only be directly connected to another element "upper" or "lower", but also indirectly connected to another element "upper" or "lower" through an intermediate element.

[0050] According to a first embodiment of the present invention, a method for generating a component tree from an assembly picture instruction based on a large language model is provided. Figure 1 The method for generating a component tree from an assembly picture description based on a large language model includes the following steps S1-S3.

[0051] Step S1: Please refer to Figure 2 and Figure 3 , identify and separate non-component legends in assembly picture instructions, and retain component information 50.

[0052] In this step, the assembly picture description generally includes component information 50 and non-component legends. In order to facilitate the subsequent generation of the component tree, it is necessary to identify and separate the non-component legends in the assembly picture description, so that the information retained in the assembly picture description is the component information.

[0053] The non-component information generally includes regular legends with regular geometric shapes as boundaries and irregular legends with irregular shapes. The irregular legends generally include auxiliary schematic legends 32 and step sequence number legends 36 that are irregular in shape but clear in meaning.

[0054] The regular legends with regular geometric shapes as borders generally include rectangular frame legends 12 with rectangles as borders and circular frame legends 16 with circles as borders. For example, the rectangular frame legend 12 can represent component details, and the running frame legend can represent connection point details.

[0055] For regular legends, traditional computer vision technology combined with segmentation methods can be used for identification and separation. It can be understood that regular legends are approximately regular geometric shapes as boundaries, which are usually not completely regular shapes and usually contain certain details, but still have regular characteristics.

[0056] Correspondingly, step S1 may include:

[0057] Regular legends with regular geometric shapes as boundaries are identified through computer vision methods.

[0058] See also Figure 2 and Figure 3 , the “identifying a regular legend with a regular geometric shape as a boundary by a computer vision method” in step S1 may specifically include:

[0059] Identify a circular frame with a circle as a boundary by using the Hough circle detection method (Example 16); and / or,

[0060] The Hough line detection method is used to identify the straight lines and screen the straight line angles. The distance between the horizontal and vertical straight line breakpoints is calculated to determine whether the two intersect and infer whether the two belong to the same rectangular frame. Based on this, a rectangular frame with a rectangle as the boundary is defined (see Figure 12).

[0061] Specifically, Hough detection is to find parts that may be straight lines or circles through the conversion between rectangular coordinate systems and coordinate systems.

[0062] For the circular frame example 16, the Hough circle detection method is used for coarse labeling.

[0063] The Hough circle is fitted by the following formula 1, where (x, y) represents a point in the image space, (a, b) represents the coordinates of the center of the circle, and R represents the radius of the circle. A circle is determined by the same a, b, and R:

[0064] Formula 1: (xa) 2 +(yb) 2 =R 2 .

[0065] The Hough circle transform transforms a circle in image space into a point in parameter space.

[0066] In a specific implementation manner, the specific process of the Hough circle detection method is as follows:

[0067] Edge detection: First, perform edge detection on the image, usually using the Canny edge detection algorithm to obtain the edge point set of the image;

[0068] Gradient calculation: Use the Sobel operator to calculate the gradient direction of each edge point. The gradient direction represents the tangent direction of the circle where the edge point is located.

[0069] Parameter space voting: traverse all possible circle center coordinates (x, y), and for each edge point, calculate its distance to the center of the circle R. If the value of the point in the accumulator of the parameter space (a, b, R) reaches the preset threshold, a circle is considered to be detected.

[0070] The parameter settings in the Hough circle detection method are set according to the assembly image size.

[0071] For the rectangular box example 12, the Hough line detection method is used for preliminary labeling.

[0072] The Hough line detection method is fitted by the following formula 2, where x and y are coordinates in the rectangular coordinate system, θ is the angle between the line and the positive direction of the x-axis, and r is the distance from the origin to the line. In theory, all points on a straight line have the same r and θ in the polar coordinate system.

[0073] Formula 2: r=xcosθ+ysinθ.

[0074] The basic principle of Hough line detection is to use the duality of points and lines, that is, the straight lines in the image space and the points in the parameter space are one-to-one corresponding. Therefore, the line detection problem in the image space is converted to the point detection problem in the parameter space, and the line detection task is completed by finding the peak in the parameter space.

[0075] In a specific implementation, the Hough line detection step may include:

[0076] Discretize θ;

[0077] Find r based on x, y and θ in the image coordinates;

[0078] Count the number of times (r, θ) appears in the offline parameter space, and the largest number is the line to be detected.

[0079] Since the component information 50 in the assembly picture manual may have perspective deformation, and the rectangular frame example 12 is usually not affected by perspective deformation, the incompletely closed rectangular frame framed by the line segments can be retrieved accordingly.

[0080] The step S1 of “identifying a regular legend with a regular geometric shape as a boundary by a computer vision method” may further include:

[0081] The recognized rule legend is compared with the mask generated by the segmentation method, and the mask whose overlap rate reaches the set value is determined as the final legend recognition result.

[0082] In this step, the preliminary result of Hough detection recognition is compared with the mask generated by the segmentation method, and the mask with a higher overlap rate is selected as the final legend recognition result.

[0083] The segmentation method for generating a mask is a technique in computer vision that is used to accurately separate objects in an image from the background; by classifying and labeling each pixel, a fine-grained division of the image area is achieved; each pixel is assigned a label to indicate whether it belongs to the foreground or background, or to a different object category; such label information forms a two-dimensional matrix, namely the segmentation mask. The generation of the segmentation mask relies on the technical support of deep learning and artificial intelligence. The neural network is trained with a large amount of image data so that it can learn and understand the characteristics of various objects. During the calculation process, the neural network gradually extracts the feature information of the image by performing operations such as convolution and pooling on the input image. Ultimately, the mask output by the network can accurately describe the position and boundaries of different objects in the image. In this embodiment, the segmentation mask can be generated and stored in the system by pre-prediction.

[0084] In this way, by fully combining the stability of traditional computer vision methods and the accuracy of segmentation algorithms, a reliable foundation is provided for subsequent legend analysis.

[0085] See also Figure 2 and Figure 3 , for the auxiliary schematic diagram legend 32 and the step number legend 36 in the irregular legend, they can be separated by combining detection with a segmentation model.

[0086] Correspondingly, step S1 may include:

[0087] An auxiliary schematic diagram example 32 with irregular shapes but clear meaning is framed by a target detection model (Gdino); and / or,

[0088] The step number legend 36 is identified by an optical character recognition method.

[0089] In this step, the auxiliary schematic diagram 32 at least includes arrows and / or hand-shaped indication images appearing in the assembly picture, which are used to indicate the operation of the assembler. The step sequence number legend 36 is generally expressed in digital form to indicate the operation sequence.

[0090] In order to achieve effective extraction of irregular legends, this embodiment adopts differentiated processing methods for different contents.

[0091] For step number legend 36, optical character recognition (OCR) technology is used for recognition to ensure that digital information can be accurately extracted.

[0092] The step of “separating the non-component legends in the assembly picture instructions” in step S1 may specifically include the following steps:

[0093] The specific coverage of non-component legends is extracted and separated through the image segmentation model, retaining the component information 50.

[0094] Image segmentation models for legend separation is a common computer vision task, which aims to identify and separate specific legends or elements from an image.

[0095] In a specific implementation, the step of extracting the specific coverage of the non-component legend by using the image segmentation model and separating it may specifically include:

[0096] Data preparation:

[0097] Collect data: A large number of annotated data sets are required, which contain legends and their corresponding labels;

[0098] Preprocessing: standardize the image, such as resizing, normalization, etc.

[0099] Model selection:

[0100] Commonly used segmentation models include:

[0101] U-Net: A classic convolutional neural network (CNN), especially suitable for medical image segmentation;

[0102] Mask R-CNN: adds a mask branch based on Faster R-CNN, which can perform object detection and instance segmentation simultaneously;

[0103] DeepLab: Improve segmentation accuracy using dilated convolution and spatial pyramid pooling modules;

[0104] Model training:

[0105] Define loss function: Common loss functions include cross entropy loss, Dice loss, etc.

[0106] Optimizer selection: The commonly used one is Adam optimizer;

[0107] Training process: Use the training data set to train the model and evaluate the performance on the validation set;

[0108] Model Evaluation:

[0109] Evaluation indicators: Commonly used evaluation indicators include IoU (Intersection over Union), Dice coefficient, Precision, Recall, etc.

[0110] Visualization results: Visualize the segmentation results to intuitively evaluate the model performance;

[0111] Model deployment:

[0112] Export model: export the trained model to a deployable format, such as ONNX, TensorFlow SavedModel, etc.

[0113] Reasoning: Using the model in real-world applications for legend separation.

[0114] For example, for auxiliary schematic diagrams 32 in assembly drawings that are irregular in shape but clear in meaning, they are separated based on text prompts through a method that combines detection and segmentation models, and the remaining parts are directly identified as component information required in the assembly steps 50. This method can achieve efficient separation and recognition of irregular shape diagrams, providing reliable data support for assembly tasks.

[0115] Step S20: Please refer to Figure 2 and Figure 4 ,The large language model organizes the assembly pictures into a step tree based on the ,few sample prompts, where the step diagram displays the ,assembly steps through a tree structure.

[0116] In this step, the task requirements and assembly pictures to be sorted are input into the large language model. Since it is still difficult for the large language model to understand the complex picture instructions, one or a small number of examples are labeled by using the few-sample prompt learning method, and the large model is required to sort the dependencies of the input steps according to the given examples. Specifically, the direct successor steps of each step are given. Through this method, the instructions of the linear assembly sequence are converted into a tree structure to display the step tree of the assembly steps, which more directly represents the dependency relationship between the steps.

[0117] More preferably, step S20 may specifically include the following steps:

[0118] Using the multimodal large model GPT-4o and the few-shot learning method, the input assembly images are sorted according to the given step tree examples to generate a step tree.

[0119] To generate the step tree, the few-shot learning method is used, combined with the multimodal large model GPT-4o, to achieve efficient and automated dependency modeling. In this process, by inputting the labeled step tree examples and the assembly step diagram to be processed into the model, the model can understand the dependencies between the steps and automatically generate a complete step tree, which clearly shows the operation sequence and logical relationship in the assembly process in the form of a tree structure.

[0120] Specifically, by utilizing the few-shot learning capability of the multimodal model, the order and dependency of each step in the assembly drawing can be identified with only a small number of labeled examples. By analyzing the input assembly step diagram, each step is accurately associated, and the output result is presented in the form of a step tree, providing an intuitive sequential structure and operation basis for the assembly process. Combining the strong reasoning ability of the large model with the high efficiency of few-shot learning, the accuracy and automation level of step dependency extraction are significantly improved, laying a solid foundation for subsequent legend information extraction and component tree construction.

[0121] Step S30: Please refer to Figure 2 and Figure 5 , combining the component information 50 and the step tree, a component tree is generated, wherein the component tree displays the premise and successor relationships between components through a tree structure.

[0122] In this step, after completing the legend processing, it is necessary to combine the information in the step tree with the result of the legend extraction to generate a complete component tree. For example, by converting the step picture of each step into the component picture without the legend, integrating the step tree, and obtaining the dependency relationship of each component in the assembly process, the component information 50 used in each step in the manual and the interdependence between the steps can be effectively extracted, which is of great significance to machine assembly.

[0123] The root node of the component tree represents the overall structure of the completed assembly. The subtrees represent each "super-part", and the tree structure intuitively displays the premise and successor relationships between components. The component tree represents: at least two premise components are assembled in sequence to form successor components until the overall structure is assembled. This hierarchical representation provides a clear framework for the modeling and management of complex assemblies. In the process of building the component tree, it is necessary to comprehensively consider the front and back dependencies of the components, and make full use of the step tree information generated from the large model to ensure the accuracy and completeness of the component tree.

[0124] Based on the dependencies in the step tree, the component tree can break down the assembly task into a logically rigorous hierarchical structure, thereby effectively supporting the modeling and optimization of the assembly process. This method not only achieves accurate modeling of complex assemblies, but also greatly improves the automation and efficiency of the assembly process, providing strong support for the efficient execution of assembly tasks.

[0125] It should be noted that the present application generates a component tree by combining the component information 50 in step S1 and the step tree in step S2. Figure 2 As shown, step S1 may be performed first, and then step S2. However, in other embodiments, step S1 and step S2 may be performed simultaneously, or step S2 may be performed first and then step S2.

[0126] In the present invention, component information is retained by identifying and separating non-component legends in assembly picture instructions. The large language model organizes the assembly picture into a step tree based on a small number of sample prompts. The component tree can be generated by combining the component information and the step tree. This effectively solves the difficulty of extracting components used for understanding the assembly sequence in the assembly picture instructions, effectively extracts the component information used in each step in the instructions, and the interdependence between the steps, which is of great significance to machine assembly.

[0127] By using traditional computer vision methods, detection models, and optical character recognition methods to jointly detect various forms of legends, and using segmentation models to remove non-component legends and retain component information, the problem of legend diversity is solved and the component information in the picture description is highlighted.

[0128] The tree-like sequential relationship between the steps generated by the few-shot learning of the large language model is combined to finally obtain the component tree. By utilizing the prior knowledge and understanding ability of the large model, the problem of difficulty in understanding the sequential relationship of complex picture instructions is solved, the accuracy of the analysis is improved, and it is helpful for subsequent machine assembly.

[0129] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0130] It should be understood that in the description of this application, unless otherwise clearly specified and limited, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance, and are not used to describe a specific order or sequence.

[0131] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0132] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be noted that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), including a number of instructions for a terminal (which can be a mobile phone, computer, control device, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0134] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for generating a component tree from an assembly picture instruction based on a large language model, characterized in that: include: Identify and separate non-component legends in assembly picture instructions and retain component information; The large language model organizes the assembly picture into a step tree based on the few sample prompts, wherein the step tree displays the assembly steps in a tree structure; The component information and the step tree are combined to generate a component tree, wherein the component tree displays the prerequisite and successor relationships between components through a tree structure.

2. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 1, characterized in that: The identifying and separating non-component legends in the assembly picture instructions and retaining component information includes: Regular legends with regular geometric shapes as boundaries are identified through computer vision methods.

3. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 2, characterized in that: After the regular legend with regular geometric shapes as boundaries is identified by computer vision methods, the following is also included: The recognized rule legend is compared with the mask generated by the segmentation method, and the mask whose overlap rate reaches the set value is determined as the final legend recognition result.

4. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 2, characterized in that: The method of identifying a regular legend with a regular geometric shape as a boundary by a computer vision method includes: Identify a circular frame with a circle as the boundary by using the Hough circle detection method; and / or, The Hough line detection method is used to identify straight lines and screen their angles. The distance between the horizontal and vertical line breakpoints is calculated to determine whether the two intersect and whether they belong to the same rectangular frame. Based on this, a rectangular frame legend with a rectangle as the boundary is defined.

5. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 4, characterized in that: The parameter settings in the Hough circle detection method are set according to the size of the assembly picture.

6. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 1, characterized in that: The identifying and separating non-component legends in the assembly picture instructions and retaining component information includes: An auxiliary schematic diagram legend with irregular shape but clear meaning is framed by a target detection model; and / or, a step sequence number legend is identified by an optical character recognition method.

7. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 6, characterized in that: The auxiliary schematic diagram example at least includes arrows and / or hand-shaped indication images appearing in the assembly picture, which are used to indicate the assembler's operation.

8. The method for generating a component tree from an assembly picture instruction based on a large language model according to any one of claims 1 to 7, characterized in that: The specific coverage of the non-component legend is extracted and separated through an image segmentation model, and the component information is retained.

9. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 1, characterized in that: The large language model organizes the assembly pictures into a step tree based on the few sample prompts, including: Using the multimodal large model GPT-4o and the few-shot learning method, the input assembly images are sorted according to the given step tree examples to generate a step tree.

10. The method for generating a component tree from an assembly picture instruction based on a large language model according to claim 1, characterized in that: The root node of the component tree represents the overall structure of the completed assembly; The component tree indicates that at least two prerequisite components are assembled in sequence to form subsequent components until the overall structure is assembled.

Citation Information

Cited By

  • Model training method and device for intelligent perception of body, and image classification method and device

    CN122289850A