Method, device, equipment, medium and program product for generating three-dimensional digital model files
By using text description information and 3D CAD modeling language data sequences, combined with pre-fine-tuned CAD modeling language generation models, 3D digital model files are automatically generated, solving the problem of low manual operation efficiency in existing technologies and improving generation quality and efficiency.
Patent Information
- Application Number
- CN202411897874.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-12-23
AI Technical Summary
Existing computer-aided design software requires a lot of manpower to generate three-dimensional digital model files, resulting in low efficiency and poor file quality.
Through text description information, the first three-dimensional digital model file and the edited three-dimensional CAD modeling language data sequence, combined with the pre-fine-tuned CAD sequence modeling language generation model, the target three-dimensional digital model file is automatically generated, reducing dependence on computer-aided design software.
It improves the quality and efficiency of generating three-dimensional digital model files, reduces labor costs, and achieves faster and more accurate digital model file determination.
Smart Images

Figure CN119337539B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and program product for generating a three-dimensional digital model file. Background Art
[0002] Computer-aided design (CAD) is a fundamental technology in industrial design. Even with the availability of CAD software, industrial design still requires extensive manual labor from skilled designers, resulting in high labor costs, low design efficiency, and poor quality of the generated 3D digital models. Summary of the Invention
[0003] The present application provides a method, apparatus, device, medium and program product for generating a three-dimensional digital model file, which can improve the quality and efficiency of generating a three-dimensional digital model file while reducing labor costs.
[0004] In a first aspect, an embodiment of the present application provides a method for generating a three-dimensional digital model file, comprising: determining a target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is the description information of the target three-dimensional digital model file or the modification parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters.
[0005] In the second aspect, an embodiment of the present application also provides a device for generating a three-dimensional digital model file, including: a target three-dimensional digital model file determination module, used to determine the target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein, the text description information is the description information of the target three-dimensional digital model file or the modification parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters.
[0006] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a three-dimensional digital model file as described in the embodiment of the present application.
[0007] In a fourth aspect, an embodiment of the present application further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the method for generating a three-dimensional digital model file as described in an embodiment of the present application.
[0008] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating a three-dimensional digital model file as described in the embodiment of the present application.
[0009] The technical solution of the embodiment of the present application determines the target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is the description information of the target three-dimensional digital model file or the modified parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, obtained in combination with editing parameters. The embodiment of the present disclosure determines the target three-dimensional digital model file by at least one of the text description information, the first three-dimensional digital model file, and the edited three-dimensional CAD modeling language data sequence, which can automatically generate the target digital model file, and can determine the target three-dimensional digital model file faster and more accurately, thereby improving the quality and efficiency of the three-dimensional digital model file generation, while reducing labor costs and reducing dependence on computer-aided design software. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 A schematic flow chart of a method for generating a three-dimensional digital model file provided in an embodiment of the present application;
[0012] Figure 2 A schematic diagram of the display effect of a first three-dimensional digital model file provided in an embodiment of the present invention;
[0013] Figure 3 A schematic diagram of the display effect of a target three-dimensional digital model file provided by an embodiment of the present invention;
[0014] Figure 4 This is a schematic diagram of the display effect of a second three-dimensional digital model file provided by an embodiment of the present invention;
[0015] Figure 5A schematic diagram of the effect of a 3D CAD modeling language data sequence provided by an embodiment of the present invention;
[0016] Figure 6 A schematic diagram of the structure of a device for generating a three-dimensional digital model file provided in an embodiment of the present application;
[0017] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should be understood that the various steps described in the method implementation of the present disclosure can be performed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. The term "including" and its variations used herein are open inclusions, that is, "including but not limited to". It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more". It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and relevant provisions.
[0020] The technical solution provided in this embodiment can generate three-dimensional digital model files in the field of three-dimensional digital modeling in the industrial field.
[0021] Figure 1 This is a flow chart of a method for generating a three-dimensional digital model file provided in an embodiment of the present application. The method can be executed by a device for generating a three-dimensional digital model file, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server. Figure 1 As shown, the method includes:
[0022] S110. Determine a target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence.
[0023] The text description information is the description information of the target three-dimensional digital model file or the modification parameter description information of the first three-dimensional digital model file.
[0024] In this embodiment, there is no restriction on the description information of the target three-dimensional digital model file. In different scenarios, the content included in the text description information may be different. For example, if the text description information is used alone (not in conjunction with the first three-dimensional digital model file) and is not used as input for the three-dimensional CAD large model intelligent entity, then the text description information is the detailed description information of the target three-dimensional digital model file. The text description information may include parameter information of the target three-dimensional digital model file (such as thickness, length, etc.), meta information (such as the name of the target three-dimensional digital model file, its use in industry, etc.), and prompt words (including format specifications and relevant content specifications of the target three-dimensional digital model file). If the text description information is used in conjunction with the first three-dimensional digital model file, then the text description information is the modified parameter description information of the first three-dimensional digital model file, such as "change the length 5 to 10".
[0025] It should be noted that, if used as input for the 3D CAD Large Model Agent in subsequent embodiments, the textual description information can be a simple description of the target 3D digital model file (including key content and requirements). This simple description is relatively concise compared to the detailed description. The 3D CAD Large Model Agent can then autonomously understand and refine the user's intent based on this textual description information through the planning module. For example, in this scenario, the textual description information could be something like "Help me generate a specific component," "Help me modify the following digital model file," or "Help me generate a 3D CAD modeling language data sequence."
[0026] The edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters.
[0027] The 3D CAD modeling language data sequence can be understood as a series of commands and parameters used to describe the process of constructing a 3D digital model file, and can be understood and executed by computer-aided design software. In this embodiment, the 3D CAD modeling language data sequence or the edited 3D CAD modeling language data sequence can be directly converted into a corresponding 3D digital model file using the computer-aided design software.
[0028] In this embodiment, the editing parameters are not limited and may be edited length, edited thickness, etc. The editing parameters may be directly applied to the 3D CAD modeling language data sequence, i.e., the editing parameters may be understood as the edited parameters in the 3D CAD modeling language data sequence.
[0029] The target 3D digital model file can be understood as a 3D digital model file that meets the user's business needs or is ultimately obtained. The first 3D digital model file can be understood as an existing, arbitrary 3D digital model file that needs to be modified or updated.
[0030] In this embodiment, the target 3D digital model file can be determined based on the text description information. For example, any large language model can be directly used to output the target 3D digital model file corresponding to the text description information based on the text description information.
[0031] In this embodiment, the target 3D digital model file can be determined based on the edited 3D CAD modeling language data sequence, that is, the edited 3D CAD modeling language data sequence can be directly converted into the target 3D digital model file using computer-aided design software.
[0032] In this embodiment, a target 3D digital model file can be determined based on the text description information and the first 3D digital model file. For example, an arbitrary large language model can be directly utilized to output a new first 3D digital model file based on the text description information and the first 3D digital model file as the target 3D digital model file. The new first 3D digital model file can be understood as a 3D digital model file that has been modified or updated based on the text description information.
[0033] In this embodiment, a target 3D digital model file can be determined based on the first 3D digital model file and the edited 3D CAD modeling language data sequence. For example, an arbitrary large language model can be used to generate a corresponding 3D CAD modeling language data sequence based on the first 3D digital model file. The corresponding 3D CAD modeling language data sequence can then be edited using editing parameters to obtain an edited 3D CAD modeling language data sequence. The edited 3D CAD modeling language data sequence can then be directly converted into a new first 3D digital model file using computer-aided design software as the target 3D digital model file.
[0034] In this embodiment, a target 3D digital model file can be determined based on the text description information, the first 3D digital model file, and the edited 3D CAD modeling language data sequence. For example, a second 3D digital model file can be output based on the text description information and / or the first 3D digital model file using any large language model. The second 3D digital model file can be a digital model file obtained based on the text description information, or a new first 3D digital model file obtained based on the text description information and the first 3D digital model file. The corresponding 3D CAD modeling language data sequence can then be generated based on the second 3D digital model file using any large language model. The corresponding 3D CAD modeling language data sequence can then be edited using editing parameters to obtain an edited 3D CAD modeling language data sequence. Computer-aided design software can then be used to directly convert the edited 3D CAD modeling language data sequence into a new second 3D digital model file, which serves as the target 3D digital model file.
[0035] The technical solution of the embodiment of the present application determines the target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is the description information of the target three-dimensional digital model file or the modified parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, obtained in combination with editing parameters. The embodiment of the present disclosure determines the target three-dimensional digital model file by at least one of the text description information, the first three-dimensional digital model file, and the edited three-dimensional CAD modeling language data sequence, which can automatically generate the target digital model file, and can determine the target three-dimensional digital model file faster and more accurately, thereby improving the quality and efficiency of the three-dimensional digital model file generation, while reducing labor costs and reducing dependence on computer-aided design software.
[0036] Optionally, determining the target three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file includes: digitally encoding the text description information to obtain a text description code; digitally encoding the first three-dimensional digital model file to obtain a first digital model code; inputting the text description code and / or the first digital model code into a pre-fine-tuned first large language model to output a second digital model code; and decoding the second digital model code to obtain the target three-dimensional digital model file.
[0037] In this embodiment, there are no restrictions on the digital encoding of text description information. For example, the text description information can be digitally encoded using the following encoding modules: a Bit-Pair encoding module, a WordPiece encoding module, a Unigram LM encoding module, and a SentencePiece encoding module. The Bit-Pair encoding module can be a Byte Pair Encoding (BPE) module, an algorithm used for text processing.
[0038] For example, the method for numerically encoding text description information is to use a tokenizer to break the text description information into discrete sub-words (tokens). Then, the corresponding serial numbers are found in a pre-set vocabulary based on the sub-words, and the entire text description information is converted into serial numbers. The numerical expression of the above text description information is projected into a text description code using an embedding network. The specific process is as follows:
[0039]
[0040] in, , They represent the corresponding sequence number and word segmenter in the vocabulary respectively. and They are subwords in the vocabulary and numerical text expressions. Represents text description information.
[0041]
[0042] Among them, Embedding means embedding network, Indicates the text description encoding.
[0043] In this embodiment, there is no restriction on the method of digitally encoding the first three-dimensional digital model file. For example, the first three-dimensional digital model file can be split into scattered faces and edges, and the faces and edges in the first three-dimensional digital model file can be encoded by a variational autoencoder and a residual vector quantization network to obtain a first digital model code.
[0044] For example, Figure 2 This is a schematic diagram of the display effect of the first three-dimensional digital model file provided by an embodiment of the present invention. The first three-dimensional digital model file takes a rounded L-shaped bracket with a single-sided hole as an example. Figure 3 A schematic diagram of the display effect of the target three-dimensional digital model file provided by an embodiment of the present invention. Figure 3 is Figure 2 On the basis of the above, a schematic diagram showing the length and height of the un-dug surface is added.
[0045] Optionally, the first three-dimensional digital model file is digitally encoded to obtain a first digital model code, including: using a first variational autoencoder to extract multiple surface features and multiple surface position features of the first three-dimensional digital model file; using a second variational autoencoder to extract multiple edge features and multiple edge position features of the first three-dimensional digital model file; fusing the multiple surface features, the multiple surface position features, the multiple edge features and the multiple edge position features to obtain fused features; encoding the fused features through a residual vector quantization network to obtain a first digital model code.
[0046] The first variational autoencoder may be a variational autoencoder for a face. The second variational autoencoder may be a variational autoencoder for an edge. The face position feature may be the center point of the corresponding face. The edge position feature may be the center point of the corresponding edge.
[0047] In this embodiment, a first variational autoencoder is used to extract multiple surface features and multiple surface position features of the first three-dimensional digital model file; a second variational autoencoder is used to extract multiple edge features and multiple edge position features of the first three-dimensional digital model file; the multiple surface features, the multiple surface position features, the multiple edge features and the multiple edge position features are fused (i.e., combined) to obtain fused features; in order to obtain more efficient coding expression, the fused features are quantized and encoded through a residual vector quantization network to obtain a first digital model code.
[0048] Exemplarily, the formula for the first digital-analog encoding may be:
[0049]
[0050] in, represents the first digital-analog code, Indicates the first to last face feature code or edge feature code or face position feature code or edge position feature code, such as Represents the first face feature encoding after the residual vector quantization network, Represents the position feature encoding of the first face after the residual vector quantization network. The encoding of the fused features, i.e. the first digital-analog encoding. Represents fusion features. N represents the total number of face feature codes, edge feature codes, face position feature codes, or edge position feature codes.
[0051] In this embodiment, the text description code and / or the first digital model code are input into the first large language model that has been fine-tuned in advance. After the second digital model code is output, that is, after all surface feature codes, all edge feature codes, all surface position feature codes, and all edge position feature codes are output, the residual vector quantization network can be used in sequence to decode all surface feature codes, all edge feature codes, all surface position feature codes, and all edge position feature codes to obtain all surface features, all edge features, all surface position features, and all edge position features. The decoder corresponding to the first variational self-encoder is then used to decode all surface features and all surface position features. The decoder corresponding to the second variational self-encoder is used to decode all edge features and all edge position features. All surfaces and all edges are then spliced according to the positions of all surfaces and all edges, thereby obtaining a target three-dimensional digital model file represented by the boundary. The boundary representation can be understood as the representation of the surface and the edge. The target three-dimensional digital model file represented by the boundary has no precision loss and can be directly used in industrial production and manufacturing.
[0052] In this embodiment, by digitally encoding the text description information and the first three-dimensional digital model file, determining the second digital model code based on the text description code and the first digital model code, and decoding the second digital model code, the generation of the target three-dimensional digital model file can be controlled in a finer granularity, thereby improving the efficiency and quality of the generation of the target three-dimensional digital model file.
[0053] Optionally, the text description code and the first digital-analog code are input into a pre-fine-tuned first large language model, and a second digital-analog code is output, including: the first large language model, based on the text description code, uses an autoregressive method to start from a starting identifier, and gradually infers the next face feature code or the next edge feature code or the next face position feature code or the next edge position feature code until the end identifier is output and the output is stopped.
[0054] In this embodiment, the first large language model uses an autoregressive approach to reason based on the text description code: starting from a start identifier, the model gradually infers the next face feature code, edge feature code, face position feature code, or edge position feature code corresponding to the target three-dimensional digital model file until an end identifier is output, at which point the output stops, thereby obtaining the inferred second digital model code. That is, each inference outputs only one face feature code, edge feature code, face position feature code, or edge position feature code (inference begins with the first face feature code, edge feature code, face position feature code, or edge position feature code, and continues until inference is complete for the last face feature code, edge feature code, face position feature code, or edge position feature code). However, each inference considers all previously output content. For example, if the first face feature code is output for the first time, the first face position feature code will be inferred based on the first face feature code during the second output. The first edge feature code will be inferred based on the first face feature code and the first face position feature code during the third output. This continues in this manner until inference is complete for the last face feature code, edge feature code, face position feature code, or edge position feature code. The end marker is the generation end marker output by the first language model, indicating that all faces, edges, edge positions, and face positions have been inferred. The start marker can be understood as the content in the text description information, which is used to remind the first language model to start inference.
[0055] In this embodiment, an autoregressive method is used to infer the feature code of each face, each edge, the position feature code of each face, and the position feature code of each edge corresponding to the target three-dimensional digital model file, which can ensure the consistency and logic of the generation process, so that the target three-dimensional digital model file finally output can accurately reflect the input text description information.
[0056] Optionally, the fine-tuning data set of the first large language model is composed of multiple first text-mathematical-model pairs, and each of the first text-mathematical-model pairs is generated as follows: parsing the first fine-tuning three-dimensional mathematic-model file to obtain three-dimensional mathematic-model parameter information; inputting the three-dimensional mathematic-model parameter information, three-dimensional mathematic-model element information and first setting prompt word corresponding to the first fine-tuning three-dimensional mathematic-model file into the second large language model, and outputting the first fine-tuning text description information; and combining the first fine-tuning text description information and the first fine-tuning three-dimensional mathematic-model file into the first text-mathematic-model pair.
[0057] The first text-to-digital model pair can be expressed as (<first fine-tuning text description information>, <first fine-tuning 3D digital model file>). The first fine-tuning 3D digital model file can be understood as any 3D digital model file used to fine-tune the first large language model.
[0058] In this embodiment, the first fine-tuned 3D digital model file can be parsed using any existing 3D digital model parsing tool (e.g., CAD software such as OpenCASCADE, CATIA, SolidWorks, AutoDesk, NX, or a deep learning-based digital model parsing module such as a neural network or Transformer network) to obtain 3D digital model parameter information (e.g., thickness, length, etc. of the first fine-tuned 3D digital model file). 3D digital model element information corresponding to the first fine-tuned 3D digital model file (e.g., the name of the first fine-tuned 3D digital model file, its industrial use, etc.) and a first setting prompt word corresponding to the first fine-tuned 3D digital model file (including formatting requirements and content requirements related to the first fine-tuned 3D digital model file) are obtained. The 3D digital model parameter information, 3D digital model element information, and first setting prompt word corresponding to the first fine-tuned 3D digital model file are then input into a second large language model, outputting a complete text description of the first fine-tuned 3D digital model file as the first fine-tuned text description information. In this embodiment, neither the second large language model nor the first large language model is limited and can be any general large language model. The second largest language model is used to expand the fine-tuning data set of the first largest language model, and the fine-tuned first largest language model is used to determine the target three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file.
[0059] In this embodiment, the fine-tuning dataset of the first large language model is generated by using the second large language model, which can improve the generation speed and quality of the fine-tuning dataset of the first large language model.
[0060] Optionally, the fine-tuning data set of the first large language model consists of multiple second text-mathematical-model pairs, and each of the second text-mathematical-model pairs is generated as follows: determining a third fine-tuning three-dimensional mathematic-model file based on the second fine-tuning three-dimensional mathematic-model file and the modification parameter description information of the second fine-tuning three-dimensional mathematic-model file; and combining the modification parameter description information of the second fine-tuning three-dimensional mathematic-model file, the second fine-tuning three-dimensional mathematic-model file, and the third fine-tuning three-dimensional mathematic-model file into the second text-mathematic-model pair.
[0061] The second text-to-digital model pair can be expressed as (<modified parameter description information of the second fine-tuned 3D digital model file, second fine-tuned 3D digital model file>-<third fine-tuned 3D digital model file>). The second fine-tuned 3D digital model file can be understood as an arbitrary 3D digital model file used to fine-tune the first language model.
[0062] In this embodiment, there is no restriction on the modification parameter description information of the second fine-tuning three-dimensional digital model file, for example
[0063] "Change the length from 5 to 10".
[0064] In this embodiment, by calling an industrial-grade three-dimensional digital model design software interface (which may be the same as a three-dimensional digital model analysis tool, such as CAD software such as CATIA and Solidworks), a new second fine-tuning three-dimensional digital model file can be automatically generated as a third fine-tuning three-dimensional digital model file based on the second fine-tuning three-dimensional digital model file and the modified parameter description information of the second fine-tuning three-dimensional digital model file (such as "changing the length 5 to 10"), so that the fine-tuning data set of the first large language model is more accurate.
[0065] Optionally, determining a target three-dimensional digital model file based on text description information and / or a first three-dimensional digital model file and an edited three-dimensional CAD modeling language data sequence includes: determining a second three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file; generating a model through the pre-fine-tuned CAD sequence modeling language, and determining the three-dimensional CAD modeling language data sequence based on the second three-dimensional digital model file or the first three-dimensional digital model file; generating an edited three-dimensional CAD modeling language data sequence based on the three-dimensional CAD modeling language data sequence in combination with the editing parameters; and converting the edited three-dimensional CAD modeling language data sequence into the target three-dimensional digital model file.
[0066] In this embodiment, the second 3D digital model file is determined based on the text description information and / or the first 3D digital model file. This is the same as the method for determining the target 3D digital model file based on the text description information and / or the first 3D digital model file, and will not be described in detail in this embodiment.
[0067] In this embodiment, a 3D CAD modeling language data sequence is generated based on the second 3D digital model file or the first 3D digital model file (or any other 3D digital model file) using the pre-adjusted CAD sequence modeling language modeling language generation model. An edited 3D CAD modeling language data sequence is generated based on the 3D CAD modeling language data sequence and the editing parameters. Computer-aided design software is used to convert the edited 3D CAD modeling language data sequence into the target 3D digital model file.
[0068] For example, Figure 4 This is a schematic diagram of the display effect of a second three-dimensional digital model file provided by an embodiment of the present invention. Figure 5 A schematic diagram of the effect of a 3D CAD modeling language data sequence provided by an embodiment of the present invention. Figure 5 is based on Figure 4 The second three-dimensional digital model file generates a three-dimensional CAD modeling language data sequence.
[0069] In this embodiment, a method of generating a three-dimensional CAD modeling language data sequence based on the pre-fine-tuned CAD sequence modeling language generation model and the second three-dimensional digital model file or the first three-dimensional digital model file can obtain a more accurate three-dimensional CAD modeling language data sequence and improve the interpretability of the target three-dimensional digital model file.
[0070] The CAD sequence modeling language generation model includes a graph neural network, a target generation model, and a decoder. In this embodiment, there is no restriction on the target generation model. The target generation model is used to generate a three-dimensional CAD sequence modeling language code corresponding to the decoder, that is, as long as the generated three-dimensional CAD sequence modeling language code can be decoded by the decoder. For example, the target generation model can be a generative variational autoencoder model, a generative transformer model, or an autoregressive generative diffusion model; it can also be a generative adversarial network (GAN), a flow-based generative model, etc.
[0071] Optionally, determining the three-dimensional CAD modeling language data sequence based on the second three-dimensional digital model file or the first three-dimensional digital model file includes: inputting the second three-dimensional digital model file or the first three-dimensional digital model file into the graph neural network, outputting connection matrix information and geometric features; inputting the connection matrix information and the geometric features into the target generation model, outputting the three-dimensional CAD sequence modeling language code; decoding the three-dimensional CAD sequence modeling language code by the decoder to obtain the three-dimensional CAD modeling language data sequence.
[0072] Among them, the connection matrix information is used to describe the topological structure of faces and edges, that is, the connection relationship between each edge and each face. The connection matrix information can be obtained through the graph topology model Characterization, where Representative surface, Represents edges. Geometric features can include surfaces for each face and curves for each edge. Surfaces and curves can be represented as regular coordinate grids and stored as node attributes (faces are equivalent to nodes) and edge attributes in the graph. Each surface can be understood as a mapping from a two-dimensional interval to a three-dimensional geometric domain.
[0073] In this embodiment, for each surface, the surface can be discretized into a regular two-dimensional sample grid, with the number of steps in the two dimensions being M and N', thereby generating M There are N' grid points, each identified by a pair of indices [k, l], where k and l are coordinates in two dimensions (for example, grid point [0, 0] is the first point in the upper left corner, and [M-1, N'-1] is the last point in the lower right corner). At each grid point indexed by [k, l], the following local features are attached: the absolute position of the 3D sampling point (that is, the specific coordinates of each grid point in 3D space). For ease of processing, the absolute positions of the 3D sampling points are normalized to a cube of size 2 centered at the origin.
[0074] The curve of each edge can define the actual geometric shape and can be understood as a mapping from a one-dimensional interval to a three-dimensional geometric domain. The curve can be a straight line, a circular arc, or a B-spline curve. In this embodiment, the curve can be discretized into a regular one-dimensional grid with M steps, that is, a one-dimensional grid consisting of M network points. At each grid point, the absolute position information of the sampling point is added as a feature, that is, each grid point also includes the specific coordinates of the point in three-dimensional space.
[0075] In this embodiment, a target generation model is used to generate a 3D CAD sequence modeling language encoding based on connection matrix information and geometric features. The 3D CAD sequence modeling language encoding is decoded by the decoder to obtain the 3D CAD modeling language data sequence. The 3D CAD modeling language data sequence includes a CAD command type, CAD command parameters, and the position of the CAD command in the 3D CAD modeling language data sequence. In this embodiment, no restrictions are placed on the CAD command type or CAD command parameters. For example, the CAD command type may be "draw edge," and the CAD command parameters may be "edge start point parameters and end point parameters."
[0076] Exemplarily, the target generation model uses an autoregressive generative diffusion model as an example. Optionally, the connection matrix information and the geometric features are input into the target generation model, and the 3D CAD sequence modeling language code is output, including: the target generation model, based on the connection matrix information and the geometric features, uses an autoregressive method to gradually generate the code of the next CAD command; each CAD command includes a CAD command type, CAD command parameters, and the position of the CAD command in the 3D CAD modeling language data sequence; the 3D CAD modeling language data sequence is composed of multiple CAD commands.
[0077] In this embodiment, the target generation model uses an autoregressive method to perform reasoning based on the connection matrix information and the geometric features: step-by-step reasoning on the code of the next CAD command corresponding to the code of the three-dimensional CAD sequence modeling language. Each reasoning only outputs the code of one CAD command (the reasoning starts from the code of the first CAD command until the reasoning of the code of the last CAD command is completed), but each reasoning will take into account all the output content. For example, the code of the first CAD command is output for the first time. When it is output for the second time, the code of the second CAD command will be inferred based on the code of the first CAD command, and so on, until the reasoning of the code of the last CAD command is completed. It should be noted that when reasoning the code of the CAD command, the target generation model starts reasoning when it encounters a start identifier and stops reasoning when it encounters an end identifier, indicating that the reasoning is complete.
[0078] For example, the coding reasoning process of multiple CAD commands can be expressed as follows:
[0079]
[0080] in, Indicates the code of the first CAD command, Indicates the code of the second CAD command. Indicates the code of the last CAD command. t indicates the total number of CAD command codes. Represents 3D CAD sequential modeling language encoding.
[0081] In this embodiment, the target generation model is used to infer the encoding of each CAD command corresponding to the 3D CAD sequence modeling language encoding in an autoregressive manner, thereby improving the accuracy of the 3D CAD sequence modeling language encoding.
[0082] Optionally, the fine-tuning data set of the CAD sequence modeling language generation model consists of multiple digital-to-analog sequence pairs; each of the digital-to-analog sequence pairs is generated as follows: through a third language model, based on the first fine-tuning three-dimensional CAD modeling language data sequence, fine-tuning editing parameters and second setting prompt words, a second fine-tuning three-dimensional CAD modeling language data sequence is generated; the digital-to-analog file corresponding to the first fine-tuning three-dimensional CAD modeling language data sequence and the second fine-tuning three-dimensional CAD modeling language data sequence are combined to form the digital-to-analog sequence pair.
[0083] Among them, the form of the digital-model sequence pair can be expressed as (<digital-model file corresponding to the first fine-tuning three-dimensional CAD modeling language data sequence>-<second fine-tuning three-dimensional CAD modeling language data sequence>). The first fine-tuning three-dimensional CAD modeling language data sequence can be an existing, arbitrary industrial three-dimensional CAD modeling language data sequence, which is used to generate more second fine-tuning three-dimensional CAD modeling language data sequences to fine-tune the CAD sequence modeling language generation model. In this embodiment, there is no restriction on the fine-tuning editing parameters, such as "change the length of 5 to 10". The second setting prompt word may include relevant format regulations and content regulations related to the second fine-tuning three-dimensional CAD modeling language data sequence. In this embodiment, there is no restriction on the third large language model, which can be any general large language model. The third large language model is used to expand the fine-tuning data set of the CAD sequence modeling language generation model.
[0084] In this embodiment, the method of determining the fine-tuning dataset of the CAD sequence modeling language generation model through the third language model can improve the generation speed and quality of the fine-tuning dataset of the CAD sequence modeling language generation model.
[0085] Optionally, determining the target three-dimensional digital model file based on the text description information includes: determining the target three-dimensional digital model file based on the text description information through a pre-fine-tuned three-dimensional CAD large model intelligent body, wherein the three-dimensional CAD large model intelligent body includes a fourth language model, a memory module and a tool module; the fourth language model includes a planning module and an execution module; wherein the memory module is used to store historical text description information; the planning module is used to determine planning instructions based on the text description information and the historical text description information; the execution module is used to call the tool module based on the planning instructions to execute the planning instructions; the tool module includes a three-dimensional digital model file generation interface, a three-dimensional digital model file editing interface and a three-dimensional CAD modeling language data sequence generation interface.
[0086] The fourth language model can be any existing large model. Historical text descriptions can be interpreted as input to the 3D CAD large model agent's history and successfully executed. The tool module provides the various tool interfaces required by the execution module, such as API calls and database access. The tool module bridges the interaction between the 3D CAD large model agent and external systems or data, supporting the agent's operational implementation and data processing.
[0087] For example, if the text description information is about generating a 3D digital model file, the planning module in the 3D CAD large model agent determines planning instructions based on the text description information. The execution module, based on the planning instructions, queries the target function in the setting data table and calls the target function, i.e., calls the 3D digital model file generation interface in the tool module. The specific process of generating the target 3D digital model file is similar to the method of "determining the target 3D digital model file based on text description information" in the aforementioned embodiment and will not be elaborated on in detail. The setting data table includes the functions of each interface in the tool module.
[0088] For example, if the text description information is about editing the content of the three-dimensional digital model file, the planning module in the three-dimensional CAD large model intelligent body determines the planning instructions based on the text description information, and the execution module queries the target function in the setting data table based on the planning instructions, and calls the target function, that is, calls the three-dimensional digital model file editing interface in the tool module. The three-dimensional digital model file editing interface specifically generates the target three-dimensional digital model file. The method of "determining the target three-dimensional digital model file based on text description information and the first three-dimensional digital model file" in the aforementioned embodiment is similar and will not be elaborated in detail.
[0089] For example, if the text description information is about generating a three-dimensional CAD modeling language data sequence, the planning module in the three-dimensional CAD large model intelligent body determines the planning instructions based on the text description information, and the execution module queries the target function in the set data table based on the planning instructions, and calls the target function, that is, calls the three-dimensional CAD modeling language data sequence generation interface in the tool module. The three-dimensional CAD modeling language data sequence generation interface specifically generates the target three-dimensional digital model file. The method of "determining the target three-dimensional digital model file based on text description information and the edited three-dimensional CAD modeling language data sequence" in the aforementioned embodiment is similar and will not be elaborated in detail.
[0090] Optionally, the memory module includes a short-term memory unit and a long-term memory unit; the short-term memory unit is used to temporarily store recent historical text description information; the long-term memory unit is used to long-term store historical text description information filtered from the recent historical text description information.
[0091] The short-term memory unit is used to temporarily store recent historical text description information. The term "recent" can be understood as historical text description information within the recent time period, such as all historical text description information input into the 3D CAD large model agent within the previous day. The short-term memory unit can also store planning instructions corresponding to the historical text description information. In this embodiment, there is no limit on the duration of the temporary period; for example, it can be one month. The short-term memory unit can be quickly accessed to respond to continuous user queries and maintain continuous knowledge of the current session. The long-term memory unit is used for long-term storage of relatively important or representative historical text description information (such as historical text description information related to frequently appearing parts) selected from the recent historical text description information. This information is persistently stored using a retrieval-augmented generation (RAG) vector library and database structure. The long-term memory unit is also used to store planning instructions corresponding to the historical text description information. In this embodiment, there is no limit on the specific duration of the long-term period; for example, it can be several years, meaning that the long-term period is greater than the temporary period. In this embodiment, RAG technology can be replaced with technologies such as Graph Retrieval-augmented generation (GraphRAG) and Knowledge Graph. The long-term memory unit can ensure that the 3D CAD large model agent can continue to evolve and adapt to the ever-changing user environment over a long period of time. In this embodiment, the database structure can be replaced by a file system, a distributed file system, an object storage, a cache system, and a log to store data.
[0092] Optionally, the planning module includes a first planning unit, a second planning unit, a third planning unit and a fourth planning unit: the first planning unit is used to determine the planning instructions based on the text description information; the second planning unit is used to determine the planning instructions based on the text description information and a third set prompt word if the execution result corresponding to the planning instructions of the first planning unit is execution failure; the third planning unit is used to determine the planning instructions based on the text description information and the historical text description information and a fourth set prompt word if the execution result corresponding to the planning instructions of the second planning unit is execution failure; the fourth planning unit is used to fine-tune the three-dimensional CAD large model intelligent body again if the execution result corresponding to the planning instructions of the third planning unit is execution failure.
[0093] The third setting prompt word, the fourth setting prompt word, and the historical text description information are used to improve the text description information to ensure the correct execution of the execution module. The planning instruction includes specific execution steps and plans.
[0094] In this embodiment, the first planning unit determines the planning instruction based on the text description information, and the execution module executes the planning instruction. During the execution process, if the execution module cannot find the objective function in the setting data table, the execution result corresponding to the planning instruction of the first planning unit is execution failure.
[0095] The second planning unit is used to determine the planning instruction based on the text description information and in combination with the third setting prompt word if the execution result corresponding to the planning instruction of the first planning unit is execution failure; the execution module executes the planning instruction. During the execution process, if the execution module cannot find the objective function in the setting data table, the execution result corresponding to the planning instruction of the second planning unit is execution failure.
[0096] The third planning unit is used to, if the execution result corresponding to the planning instruction of the second planning unit is execution failure, determine the planning instruction based on the text description information and the historical text description information, combined with the fourth setting prompt word; the execution module executes the planning instruction, and during the execution process, if the execution module cannot find the objective function in the setting data table, the execution result corresponding to the planning instruction of the third planning unit is execution failure.
[0097] The fourth planning unit is used to call the execution module again to fine-tune the three-dimensional CAD large model intelligent body if the execution result corresponding to the planning instruction of the third planning unit is execution failure.
[0098] In this embodiment, if the text description information is a description of a simple generation task, such as a simple mathematical calculation, the execution module can be directly executed according to the planning module without calling the tool module.
[0099] It should be noted that the tool module can be embedded in the computer-aided design software as a function of the computer-aided design software.
[0100] The fine-tuning dataset for the 3D CAD large model agent consists of multiple text planning instruction pairs, each of which takes the form (<fine-tuning text>, <planning instruction>). This dataset is used to fine-tune the 3D CAD large model agent. Based on the fine-tuning text and operation sequences, the agent learns planning capabilities, improves the planning module's capabilities, and makes the generated target 3D digital model file more accurate.
[0101] In this embodiment, the 3D CAD large model intelligent body can receive any text description information from the user, and can independently call the computer-aided design software for manual operation and self-editing to generate the required industrial-grade target 3D digital model file.
[0102] Figure 6A schematic diagram of a device for generating a three-dimensional digital model file provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the apparatus includes: a target three-dimensional digital model file determination module 610;
[0103] The target three-dimensional digital model file determination module 610 is used to determine the target three-dimensional digital model file based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is the description information of the target three-dimensional digital model file or the modification parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters.
[0104] The technical solution of the embodiment of the present application is to determine the target three-dimensional digital model file through the target three-dimensional digital model file determination module based on at least one of the following methods: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein, the text description information is the description information of the target three-dimensional digital model file or the modification parameter description information of the first three-dimensional digital model file; the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, obtained in combination with editing parameters. In the embodiment of the present disclosure, the method of determining the target three-dimensional digital model file through at least one of the text description information, the first three-dimensional digital model file, and the edited three-dimensional CAD modeling language data sequence can automatically generate the target digital model file, and can determine the target three-dimensional digital model file faster and more accurately, thereby improving the quality and efficiency of three-dimensional digital model file generation, while reducing labor costs and reducing dependence on computer-aided design software.
[0105] Optionally, the target three-dimensional digital model file determination module is specifically used to: digitally encode the text description information to obtain a text description code; digitally encode the first three-dimensional digital model file to obtain a first digital model code; input the text description code and / or the first digital model code into a pre-fine-tuned first large language model to output a second digital model code; decode the second digital model code to obtain the target three-dimensional digital model file.
[0106] Optionally, the target three-dimensional digital model file determination module is also used to: use a first variational autoencoder to extract multiple surface features and multiple surface position features of the first three-dimensional digital model file; use a second variational autoencoder to extract multiple edge features and multiple edge position features of the first three-dimensional digital model file; fuse the multiple surface features, the multiple surface position features, the multiple edge features and the multiple edge position features to obtain fused features; encode the fused features through a residual vector quantization network to obtain a first digital model code.
[0107] Optionally, the target three-dimensional digital model file determination module is specifically used for: the first largest language model, based on the text description code, adopts an autoregressive method starting from the starting identifier, and gradually infers the next face feature code or the next edge feature code or the next face position feature code or the next edge position feature code until the output end identifier is output and the output is stopped.
[0108] The fine-tuning data set of the first large language model is composed of multiple first text-mathematical-model pairs. Optionally, the above-mentioned device also includes a data set generation module, which is used to parse the first fine-tuning three-dimensional mathematical-model file to obtain three-dimensional mathematical-model parameter information; input the three-dimensional mathematical-model parameter information, three-dimensional mathematical-model element information and first setting prompt word corresponding to the first fine-tuning three-dimensional mathematical-model file into the second large language model, and output the first fine-tuning text description information; the first fine-tuning text description information and the first fine-tuning three-dimensional mathematical-model file are combined into the first text-mathematical-model pair.
[0109] The fine-tuning data set of the first large language model is composed of multiple second text-mathematical-model pairs. Optionally, the above-mentioned device also includes a data set generation module, which is also used to determine the third fine-tuning three-dimensional digital-model file based on the second fine-tuning three-dimensional digital-model file and the modification parameter description information of the second fine-tuning three-dimensional digital-model file; the modification parameter description information of the second fine-tuning three-dimensional digital-model file, the second fine-tuning three-dimensional digital-model file and the third fine-tuning three-dimensional digital-model file are combined into the second text-mathematical-model pair.
[0110] Optionally, the target three-dimensional digital model file determination module is also used to: determine a second three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file; generate a model through the pre-fine-tuned CAD sequence modeling language, and determine the three-dimensional CAD modeling language data sequence based on the second three-dimensional digital model file or the first three-dimensional digital model file; generate an edited three-dimensional CAD modeling language data sequence based on the three-dimensional CAD modeling language data sequence in combination with the editing parameters; and convert the edited three-dimensional CAD modeling language data sequence into the target three-dimensional digital model file.
[0111] The CAD sequential modeling language generation model includes a graph neural network, a target generation model, and a decoder. The target generation model is used to generate a three-dimensional CAD sequential modeling language code corresponding to the decoder.
[0112] Optionally, the target three-dimensional digital model file determination module is also used to: input the second three-dimensional digital model file or the first three-dimensional digital model file into the graph neural network, and output connection matrix information and geometric features; input the connection matrix information and the geometric features into the target generation model, and output the three-dimensional CAD sequence modeling language code; decode the three-dimensional CAD sequence modeling language code through the decoder to obtain the three-dimensional CAD modeling language data sequence.
[0113] The target generation model is an autoregressive generation diffusion model.
[0114] Optionally, the target three-dimensional digital model file determination module is also used for: the target generation model, based on the connection matrix information and the geometric features, adopts an autoregressive method to gradually generate the encoding of the next CAD command; each of the CAD commands includes a CAD command type, CAD command parameters and the position of the CAD command in the three-dimensional CAD modeling language data sequence; the three-dimensional CAD modeling language data sequence is composed of multiple CAD commands.
[0115] The fine-tuning dataset for the CAD sequence modeling language generation model is composed of multiple digital-to-analog sequence pairs. Optionally, the apparatus further includes a dataset generation module, further configured to generate a second fine-tuning 3D CAD modeling language data sequence using a third language model based on the first fine-tuning 3D CAD modeling language data sequence, fine-tuning editing parameters, and a second setting prompt; and to combine the digital-to-analog sequence pair with the digital-to-analog file corresponding to the first fine-tuning 3D CAD modeling language data sequence and the second fine-tuning 3D CAD modeling language data sequence.
[0116] Optionally, the target three-dimensional digital model file determination module is also used to: determine the target three-dimensional digital model file based on the text description information through a pre-fine-tuned three-dimensional CAD large model intelligent body, wherein the three-dimensional CAD large model intelligent body includes a fourth language model, a memory module and a tool module; the fourth language model includes a planning module and an execution module; wherein the memory module is used to store historical text description information; the planning module is used to determine planning instructions based on the text description information and the historical text description information; the execution module is used to call the tool module based on the planning instructions to execute the planning instructions; the tool module includes a three-dimensional digital model file generation interface, a three-dimensional digital model file editing interface and a three-dimensional CAD modeling language data sequence generation interface.
[0117] Optionally, the memory module includes a short-term memory unit and a long-term memory unit; the short-term memory unit is used to temporarily store recent historical text description information; the long-term memory unit is used to long-term store historical text description information filtered from the recent historical text description information.
[0118] Optionally, the planning module includes a first planning unit, a second planning unit, a third planning unit and a fourth planning unit: the first planning unit is used to determine the planning instructions based on the text description information; the second planning unit is used to determine the planning instructions based on the text description information and a third set prompt word if the execution result corresponding to the planning instructions of the first planning unit is execution failure; the third planning unit is used to determine the planning instructions based on the text description information and the historical text description information and a fourth set prompt word if the execution result corresponding to the planning instructions of the second planning unit is execution failure; the fourth planning unit is used to fine-tune the three-dimensional CAD large model intelligent body again if the execution result corresponding to the planning instructions of the third planning unit is execution failure.
[0119] The device for generating a three-dimensional digital model file provided in the embodiment of the present application can execute the method for generating a three-dimensional digital model file provided in any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.
[0120] Figure 7 A schematic diagram of an electronic device 10 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0121] like Figure 7As shown, electronic device 10 includes at least one processor 11 and memory, such as read-only memory (ROM) 12 and random access memory (RAM) 13, communicatively connected to at least one processor 11. The memory stores computer programs executable by the at least one processor. Processor 11 can perform various appropriate actions and processes based on the computer programs stored in ROM 12 or loaded from storage unit 18 into RAM 13. RAM 13 can also store various programs and data required for the operation of electronic device 10. Processor 11, ROM 12, and RAM 13 are interconnected via bus 14. An input / output (I / O) interface 15 is also connected to bus 14.
[0122] Various components in electronic device 10 are connected to an input / output (I / O) interface 15, including an input unit 16, such as a keyboard and mouse; an output unit 17, such as various types of displays and speakers; a storage unit 18, such as a magnetic disk and optical disk; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0123] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the method for generating a three-dimensional digital model file.
[0124] In some embodiments, the method for generating a three-dimensional digital model file can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via a read-only memory (ROM) 12 and / or a communication unit 19. When the computer program is loaded into a random access memory (RAM) 13 and executed by the processor 11, one or more steps of the method for generating a three-dimensional digital model file described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for generating a three-dimensional digital model file in any other appropriate manner (for example, by means of firmware).
[0125] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0126] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0127] In the context of the present application, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or apparatus. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0129] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0130] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0131] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating a three-dimensional digital model file as provided in any embodiment of the present application.
[0132] In the process of implementation, the computer program product can be written in one or more programming languages or a combination thereof to write computer program code for performing the operations of the present application, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0133] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for generating a three-dimensional digital model file, characterized in that: include: Determining a target three-dimensional digital model file based on at least one of the following: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is description information of the target three-dimensional digital model file or modified parameter description information of the first three-dimensional digital model file; and the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters; wherein, determining the target three-dimensional digital model file based on the text description information and the first three-dimensional digital model file includes: numerically encoding the text description information to obtain a text description code; numerically encoding the first three-dimensional digital model file to obtain a first digital model code; inputting the text description code and the first digital model code into a pre-fine-tuned first large language model to output a second digital model code; decoding the second digital model code to obtain the target three-dimensional digital model file; wherein the fine-tuning data set of the first large language model is composed of a plurality of second text digital model pairs; wherein the first digital model code is a first digital model code obtained by encoding the faces and edges in the first three-dimensional digital model file through a variational autoencoder and a residual vector quantization network; wherein the second digital model code includes all face feature codes, all edge feature codes, all face position feature codes, and all edge position feature codes corresponding to the target three-dimensional digital model file; wherein the method for decoding the second digital model code is: decoding the second digital model code through the variational autoencoder; wherein the target three-dimensional digital model file is a target three-dimensional digital model file represented by a boundary; The step of digitally encoding the first three-dimensional digital model file to obtain a first digital model code includes: Extracting multiple surface features and multiple surface position features of the first three-dimensional digital model file using a first variational autoencoder; Extracting a plurality of edge features and position features of the plurality of edges of the first three-dimensional digital model file using a second variational autoencoder; fusing the plurality of surface features, the plurality of surface position features, the plurality of edge features, and the plurality of edge position features to obtain a fused feature; Encoding the fused features through a residual vector quantization network to obtain a first digital-analog code; The text description code and the first digital-analog code are input into a pre-fine-tuned first large language model, and a second digital-analog code is output, including: The first language model, based on the text description code, uses an autoregressive method to start from a starting identifier, gradually inferring the next face feature code or the next edge feature code or the next face position feature code or the next edge position feature code until an end identifier is output and the output stops.
2. The method according to claim 1, characterized in that Determining a target three-dimensional digital model file based on the text description information or the first three-dimensional digital model file includes: Numerically encoding the text description information to obtain a text description code; Digitally encoding the first three-dimensional digital model file to obtain a first digital model code; Inputting the text description code or the first digital-analog code into a pre-fine-tuned first large language model, and outputting a second digital-analog code; The second digital model code is decoded to obtain the target three-dimensional digital model file.
3. The method according to claim 2, characterized in that The fine-tuning dataset of the first language model consists of a plurality of first text-digital-model pairs, and each of the first text-digital-model pairs is generated as follows: Parsing the first fine-tuning three-dimensional digital model file to obtain three-dimensional digital model parameter information; Inputting the three-dimensional digital model parameter information, three-dimensional digital model element information, and first setting prompt word corresponding to the first fine-tuning three-dimensional digital model file into the second language model, and outputting first fine-tuning text description information; The first fine-tuning text description information and the first fine-tuning three-dimensional digital model file are combined into the first text-digital model pair.
4. The method according to claim 1, wherein Each second text number-module pair is generated as follows: Determining a third fine-tuning three-dimensional digital model file based on the second fine-tuning three-dimensional digital model file and the modification parameter description information of the second fine-tuning three-dimensional digital model file; The modification parameter description information of the second fine-tuning three-dimensional digital model file, the second fine-tuning three-dimensional digital model file and the third fine-tuning three-dimensional digital model file are combined into the second text-digital model pair.
5. The method according to claim 1, wherein Determining a target three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file and the edited three-dimensional CAD modeling language data sequence includes: Determining a second three-dimensional digital model file based on the text description information and / or the first three-dimensional digital model file; Generate a model by using the pre-adjusted CAD sequence modeling language, and determine the 3D CAD modeling language data sequence based on the second 3D digital model file or the first 3D digital model file; generating an edited 3D CAD modeling language data sequence based on the 3D CAD modeling language data sequence and the editing parameters; The edited three-dimensional CAD modeling language data sequence is converted into the target three-dimensional digital model file.
6. The method according to claim 5, characterized in that in, The CAD sequence modeling language generation model includes a graph neural network, a target generation model, and a decoder; wherein the target generation model is used to generate a three-dimensional CAD sequence modeling language code corresponding to the decoder; determining the three-dimensional CAD modeling language data sequence based on the second three-dimensional digital model file or the first three-dimensional digital model file includes: Inputting the second three-dimensional digital model file or the first three-dimensional digital model file into the graph neural network, and outputting connection matrix information and geometric features; Inputting the connection matrix information and the geometric features into a target generation model, and outputting the 3D CAD sequential modeling language code; The three-dimensional CAD sequence modeling language code is decoded by the decoder to obtain the three-dimensional CAD modeling language data sequence.
7. The method according to claim 6, characterized in that in, The target generation model is an autoregressive generative diffusion model; the connection matrix information and the geometric features are input into the target generation model, and the 3D CAD sequential modeling language code is output, including: The target generation model uses an autoregressive method to gradually generate the encoding of the next CAD command based on the connection matrix information and the geometric features; each CAD command includes a CAD command type, CAD command parameters and the position of the CAD command in the three-dimensional CAD modeling language data sequence; the three-dimensional CAD modeling language data sequence is composed of multiple CAD commands.
8. The method according to claim 1, 5 or 6, characterized in that: The fine-tuning dataset of the CAD sequential modeling language generation model consists of multiple digital-analog sequence pairs; each of the digital-analog sequence pairs is generated as follows: Generate a second fine-tuned 3D CAD modeling language data sequence based on the first fine-tuned 3D CAD modeling language data sequence, the fine-tuned editing parameters, and the second set prompt word using the third language model; The digital-model file corresponding to the first fine-tuning 3D CAD modeling language data sequence and the second fine-tuning 3D CAD modeling language data sequence are combined into the digital-model sequence pair.
9. The method according to claim 1, characterized in that Determining a target three-dimensional digital model file based on the text description information includes: Through the pre-fine-tuned 3D CAD large model intelligent body, the target 3D digital model file is determined based on the text description information, wherein the 3D CAD large model intelligent body includes a fourth language model, a memory module and a tool module; the fourth language model includes a planning module and an execution module; wherein the memory module is used to store historical text description information; the planning module is used to determine planning instructions based on the text description information and the historical text description information; the execution module is used to call the tool module based on the planning instructions to execute the planning instructions; the tool module includes a 3D digital model file generation interface, a 3D digital model file editing interface and a 3D CAD modeling language data sequence generation interface.
10. The method according to claim 9, characterized in that The memory module includes a short-term memory unit and a long-term memory unit; the short-term memory unit is used to temporarily store recent historical text description information; the long-term memory unit is used to long-term store historical text description information filtered from the recent historical text description information.
11. The method according to claim 9, characterized in that The planning module includes a first planning unit, a second planning unit, a third planning unit and a fourth planning unit: the first planning unit is used to determine the planning instructions based on the text description information; the second planning unit is used to determine the planning instructions based on the text description information and a third set prompt word if the execution result corresponding to the planning instructions of the first planning unit is execution failure; the third planning unit is used to determine the planning instructions based on the text description information and the historical text description information and a fourth set prompt word if the execution result corresponding to the planning instructions of the second planning unit is execution failure; the fourth planning unit is used to fine-tune the three-dimensional CAD large model intelligent body again if the execution result corresponding to the planning instructions of the third planning unit is execution failure.
12. A device for generating a three-dimensional digital model file, characterized in that: include: a target three-dimensional digital model file determination module, configured to determine a target three-dimensional digital model file based on at least one of the following: text description information, a first three-dimensional digital model file, and an edited three-dimensional CAD modeling language data sequence; wherein the text description information is description information of the target three-dimensional digital model file or modified parameter description information of the first three-dimensional digital model file; and the edited three-dimensional CAD modeling language data sequence is a three-dimensional CAD modeling language data sequence generated by a pre-fine-tuned CAD sequence modeling language generation model, combined with editing parameters; wherein, the target three-dimensional digital model file determination module is used to: digitally encode the text description information to obtain a text description code; digitally encode the first three-dimensional digital model file to obtain a first digital model code; input the text description code and the first digital model code into a pre-fine-tuned first large language model, and output a second digital model code; decode the second digital model code to obtain the target three-dimensional digital model file; wherein the fine-tuning data set of the first large language model is composed of multiple second text digital model pairs; wherein the first digital model code is a first digital model code obtained by encoding the faces and edges in the first three-dimensional digital model file through a variational autoencoder and a residual vector quantization network; wherein the second digital model code includes all face feature codes, all edge feature codes, all face position feature codes, and all edge position feature codes corresponding to the target three-dimensional digital model file; wherein the method for decoding the second digital model code is: decoding the second digital model code through the variational autoencoder; wherein the target three-dimensional digital model file is a target three-dimensional digital model file represented by a boundary; The target three-dimensional digital model file determination module is further configured to: extract multiple surface features and multiple surface position features of the first three-dimensional digital model file using a first variational autoencoder; Extracting a plurality of edge features and position features of the plurality of edges of the first three-dimensional digital model file using a second variational autoencoder; fusing the plurality of surface features, the plurality of surface position features, the plurality of edge features, and the plurality of edge position features to obtain a fused feature; Encoding the fused features through a residual vector quantization network to obtain a first digital-analog code; Among them, the target three-dimensional digital model file determination module is also used for: the first large language model, based on the text description code, adopts an autoregressive method to start from the starting identifier, and gradually infers the next face feature code or the next edge feature code or the next face position feature code or the next edge position feature code until the end identifier is output and the output is stopped.
13. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a three-dimensional digital model file as described in any one of claims 1 to 11.
14. A storage medium comprising computer executable instructions, wherein the computer executable instructions, when executed by a computer processor, are used to execute the method for generating a three-dimensional digital model file according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for generating a three-dimensional digital model file according to any one of claims 1 to 11.
Citation Information
Patent Citations
Workpiece three-dimensional model design generation method and system based on large language model
CN118012416A