Method and device for realizing self-evolution of three-dimensional assembly intelligent agent by computing power through intelligent computing cloud platform

CN122595850APending Publication Date: 2026-08-18DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611008298.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0009]本发明提供一种智能计算云平台通过算力实现三维装配智能体自进化的方法及装置,以解决现有技术中三维装配自动化不稳定、装配结果不易评估、优化无法持续迭代、训练数据难以沉淀以及小模型部署效果不足的问题

Benefits of technology

[0024]In this invention, step S1 involves scheduling computing power in an intelligent computing cloud platform to generate configuration data for each business module based on the 3D files uploaded by the user for multiple business modules. These business modules include an assembly module, an inspection module, and a reinforcement learning module. Step S2 involves obtaining feature information of the parts to be assembled based on the configuration data for the assembly module, processing the feature information, and generating candidate layout schemes. Step S3 involves inspecting the candidate layout schemes based on the configuration data for the inspection module, obtaining inspection results, and updating the state parameters of the current session. Step S4 involves determining the exploration strategy and reward function value based on the inspection results, the state parameters of the current session, and the candidate layout schemes, and updating the dynamic running experience base. Step S5 involves generating new candidate layout schemes based on the exploration strategy and reward function value, obtaining the inspection results and state parameters corresponding to the new candidate layouts, and constructing preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met. Step S6 involves distilling the student model based on the candidate layout schemes, reward function values, the dynamic running experience base, and preference samples to obtain the target assembly model. In this way, the present invention can rely on the computing power scheduling and unified service capabilities of the intelligent computing cloud platform to automatically generate a layout scheme for three-dimensional assembly, forming a replicable technical path for industrial enterprises. The layout scheme is then intelligently inspected to obtain inspection results. Reinforcement learning is used to achieve online self-evolution of the assembly strategy. Furthermore, by constructing preference samples, the assembly capabilities of the large model are distilled into small and medium-sized models that are more suitable for industrial implementation. This significantly reduces the model inference cost and the engineering deployment threshold while ensuring the success rate of automatic assembly and the interpretability of inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595850A_ABST
    Figure CN122595850A_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for achieving self-evolution of a 3D assembly intelligent agent through computing power on an intelligent computing cloud platform. It relates to the fields of intelligent computing centers, smart computing centers, computing infrastructure, and smart computing cloud technologies, and includes the following steps: Step S1, generating configuration data based on a 3D file; Step S2, acquiring feature information and generating candidate layout schemes; Step S3, obtaining the inspection results of the candidate layout schemes and updating state parameters; Step S4, determining the exploration strategy and reward function value based on the inspection results, state parameters, and candidate layout schemes, and updating the experience base; Step S5, generating new candidate layout schemes based on the exploration strategy and reward function value, and constructing preference samples based on the inspection results and state parameters of all candidate layout schemes; Step S6, distilling the student model to obtain the target assembly model. This invention can significantly reduce model inference costs and engineering deployment thresholds while ensuring automatic assembly success rate and inspection interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent computing centers, smart computing centers, computing infrastructure, and smart cloud computing technologies, specifically to a method and apparatus for realizing the self-evolution of a three-dimensional assembled intelligent body through computing power on an intelligent computing cloud platform. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to execute a certain computing requirement. It is the computing power to achieve the target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] 3D assembly is a crucial step in industrial design, process verification, manufacturing preparation, and digital delivery. Currently, assembly software typically requires engineers to manually read parts, establish constraints, adjust layouts, analyze gaps, check for interferences, and repeatedly modify them within the interface of 3D engineering drawing software. For assembly objects of dozens, hundreds, or even larger scales, this manual approach is not only time-consuming but also susceptible to differences in experience, context switching, operational errors, and fluctuations in interface status, resulting in an inefficient and unstable assembly process. Faced with this insurmountable efficiency bottleneck of traditional methods, the industry urgently needs to leverage the powerful computing capabilities of intelligent computing centers to explore a paradigm shift from manual to data- and algorithm-driven approaches.

[0008] It is evident that since the emergence of intelligent computing centers, how to free 3D assembly tasks from strong reliance on manual labor, automatically evaluate assembly schemes without definite answers, realize the self-evolution of 3D assembly intelligent agents, and reduce model reasoning costs and engineering deployment thresholds has been a pressing problem in this field. Summary of the Invention

[0009] This invention provides a method and apparatus for realizing the self-evolution of a three-dimensional assembly intelligent agent through computing power on an intelligent computing cloud platform, in order to solve the problems of unstable automation of three-dimensional assembly, difficulty in evaluating assembly results, inability to continuously iterate optimization, difficulty in accumulating training data, and insufficient deployment effect of small models in the prior art.

[0010] To solve the above problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for an intelligent computing cloud platform to achieve self-evolution of a three-dimensional assembly intelligent agent through computing power, comprising: Step S1: Schedule the computing power in the intelligent computing cloud platform, and generate configuration data for each business module based on the three-dimensional files uploaded by the user for multiple business modules. The multiple business modules include an assembly module, an inspection module, and a reinforcement learning module. Step S2: Based on the configuration data corresponding to the assembly module, obtain the feature information of the part to be assembled, process the feature information, and generate a candidate layout scheme. Step S3: Based on the configuration data corresponding to the inspection module, inspect the candidate layout scheme, obtain the inspection result, and update the state parameters of the current session; Step S4: Based on the inspection results, the current session's state parameters, and the candidate layout scheme, determine the exploration strategy and reward function value, and update the dynamic running experience base; Step S5: Generate new candidate layout schemes based on the exploration strategy and the reward function value, obtain the inspection results and state parameters corresponding to the new candidate layouts, and construct preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met. Step S6: Based on the candidate layout scheme, the reward function value, the dynamic running experience library, and the preference sample, the student model is distilled to obtain the target assembly model.

[0011] In one embodiment, step S4 includes: Step S4.1: Determine the exploration strategy based on the state parameters of the current session. The state parameters include the current best score, historical score sequence, stagnation count, layout signature set, and dynamic experience base summary. Step S4.2: Determine the reward function value based on the score and number of risk penalty items in the inspection results, and the structural legality index of the candidate layout scheme; Step S4.3: Based on the inspection results, determine the set of experience items for this layout, and perform fusion calculation with the historical dynamic operation experience library to obtain the updated dynamic operation experience library.

[0012] In one embodiment, step S4.1 includes at least one of the following: Step S4.1.1: Calculate the selection probability of each candidate exploration strategy based on the state parameters of the current session, and determine the candidate exploration strategy with the highest selection probability as the target exploration strategy. The candidate exploration strategies include baseline exploration, local exploration, global exploration, and strong exploration. Step S4.1.2: Based on the current best score and the historical score sequence, calculate the stagnation index. If the stagnation index exceeds a first threshold, determine the target exploration strategy as strong exploration. Step S4.1.3: Based on the number of times the layout signature corresponding to the candidate layout scheme appears in the layout signature set, if the number of times exceeds the second threshold, determine the target exploration strategy as transforming the candidate layout scheme.

[0013] In one embodiment, after step S4.2, the method further includes: Step S4.4: In manual mode, obtain the human scores for the candidate layout schemes. Step S4.5: Based on the human scoring, the coefficients of the human scoring mapped to the unified reward space, and the reward function value, determine the reward function value after human feedback fusion.

[0014] In one embodiment, step S6 includes: Step S6.1: Based on the candidate layout scheme, the reward function value, the dynamic running experience base, and the preference sample, determine at least two distillation losses corresponding to supervised distillation, preference distillation, reward distillation, and context distillation respectively; Step S6.2: The distillation loss is weighted and summed to obtain the comprehensive loss. The student model is trained using the comprehensive loss to obtain the target assembly model.

[0015] In one embodiment, step S6.1 includes at least two of the following: Step S6.1.1: Based on the candidate layout scheme, determine the supervised distillation loss; Step S6.1.2: Determine the preference distillation loss based on the preference sample; Step S6.1.3: Determine the reward distillation loss based on the reward function value; Step S6.1.4: Determine the context distillation loss based on the dynamic operation experience base.

[0016] In one embodiment, step S2 includes: Step S2.1: Based on the current requirements of the 3D assembly scene, determine the feature information extraction mode of the parts to be assembled; Step S2.2: Based on the information range corresponding to the extraction mode, extract the feature information of the part to be assembled from the configuration data.

[0017] In one embodiment, step S2 includes: Step S2.3: Determine the number of parts to be assembled in the 3D file contained in the configuration data; Step S2.4: Determine the prompt template for the initial stage based on the relationship between the quantity and the quantity threshold; Step S2.5: Generate candidate layout schemes based on the prompt template from the initial stage and the feature information; Step S2.6: If the candidate layout scheme is non-compliant, update the candidate layout scheme based on the prompt template for the next stage and the feature information.

[0018] In one embodiment, step S2 includes: Step S2.1': Input the prompt template containing the output format and the feature information into the large model to obtain the candidate layout scheme generated by the large model. The output format indicates that the candidate layout scheme is output in JSON mode and includes at least anchor index and assembly constraint relationship.

[0019] In one embodiment, after step S2, the method further includes: Step S7: Upon receiving a download instruction, create a temporary assembly document and write the candidate layout scheme into the temporary assembly document; Step S8: If the temporary assembly document is saved to the download directory, close the temporary assembly document and restore the original active document.

[0020] Secondly, the present invention also provides a device for an intelligent computing cloud platform to achieve self-evolution of a three-dimensional assembly intelligent agent through computing power, comprising: The silent task management module is used to schedule computing power in the intelligent computing cloud platform. Based on the three-dimensional files uploaded by the user for multiple business modules, it generates configuration data corresponding to each business module. The multiple business modules include an assembly module, an inspection module, and a reinforcement learning module. An automatic assembly module is used to obtain feature information of the parts to be assembled based on the configuration data corresponding to the assembly module, process the feature information, and generate candidate layout schemes. The intelligent inspection module is used to inspect the candidate layout scheme based on the configuration data corresponding to the inspection module, obtain the inspection results, and update the status parameters of the current session. The reinforcement learning module is used to determine the exploration strategy and reward function value based on the inspection results, the state parameters of the current session, and the candidate layout scheme, and to update the dynamic running experience base. The sample determination module is used to generate new candidate layout schemes based on the exploration strategy and the reward function value, obtain the inspection results and state parameters corresponding to the new candidate layout schemes, and construct preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes when the iteration termination condition corresponding to the reinforcement learning module is met. The distillation training module is used to distill the student model based on the candidate layout scheme, the reward function value, the dynamic running experience library, and the preference sample to obtain the target assembly model.

[0021] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements the steps in the method for the self-evolution of a three-dimensional assembly intelligent agent by computing power through the intelligent computing cloud platform described in the first aspect above.

[0022] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the method for the self-evolution of a three-dimensional assembly intelligent body through computing power by an intelligent computing cloud platform as described in the first aspect above.

[0023] Fifthly, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps in the method for the self-evolution of a three-dimensional assembly intelligent agent by computing power through an intelligent computing cloud platform as described in the first aspect above.

[0024] In this invention, step S1 involves scheduling computing power in an intelligent computing cloud platform to generate configuration data for each business module based on the 3D files uploaded by the user for multiple business modules. These business modules include an assembly module, an inspection module, and a reinforcement learning module. Step S2 involves obtaining feature information of the parts to be assembled based on the configuration data for the assembly module, processing the feature information, and generating candidate layout schemes. Step S3 involves inspecting the candidate layout schemes based on the configuration data for the inspection module, obtaining inspection results, and updating the state parameters of the current session. Step S4 involves determining the exploration strategy and reward function value based on the inspection results, the state parameters of the current session, and the candidate layout schemes, and updating the dynamic running experience base. Step S5 involves generating new candidate layout schemes based on the exploration strategy and reward function value, obtaining the inspection results and state parameters corresponding to the new candidate layouts, and constructing preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met. Step S6 involves distilling the student model based on the candidate layout schemes, reward function values, the dynamic running experience base, and preference samples to obtain the target assembly model. In this way, the present invention can rely on the computing power scheduling and unified service capabilities of the intelligent computing cloud platform to automatically generate a layout scheme for three-dimensional assembly, forming a replicable technical path for industrial enterprises. The layout scheme is then intelligently inspected to obtain inspection results. Reinforcement learning is used to achieve online self-evolution of the assembly strategy. Furthermore, by constructing preference samples, the assembly capabilities of the large model are distilled into small and medium-sized models that are more suitable for industrial implementation. This significantly reduces the model inference cost and the engineering deployment threshold while ensuring the success rate of automatic assembly and the interpretability of inspection. Attached Figure Description

[0025] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart of a method for realizing the self-evolution of a three-dimensional assembly intelligent body through computing power using an intelligent computing cloud platform provided by the present invention; Figure 2 This is a structural diagram of a device provided by the present invention that enables the self-evolution of a three-dimensional assembly intelligent body through computing power in an intelligent computing cloud platform. Figure 3 This is a structural diagram of an electronic device provided by the present invention. Detailed Implementation

[0027] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0028] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0029] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0030] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0031] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices within servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), while the commonly used unit of measurement for performance is the number of read / write operations per second (IOPS / TB). Disaster recovery ratio is an important indicator of security and reliability.

[0032] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0033] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0034] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0035] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0036] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.

[0037] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0038] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0039] The "intelligent computing cloud platform" mentioned in this invention, abbreviated as "intelligent computing cloud", refers to a cloud computing platform that integrates hardware and software resources based on an intelligent computing center.

[0040] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0041] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0042] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0043] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0044] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0045] The "agent" described in this invention refers to an agent capable of perceiving the environment and taking actions to achieve specific goals. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. The agent perceives changes in the environment (e.g., through sensors or data input), makes judgments and decisions based on its learned knowledge and algorithms, and then executes actions to influence the environment or achieve predetermined goals.

[0046] The "silent task" described in this invention refers to an independent task payload formed by uploading files, which independently completes the tasks of reading, analyzing, optimizing, and exporting without relying on the current active window of the CAD graphical interface.

[0047] The "silent task isolation" described in this invention refers to a mechanism in which the three modules of assembly, inspection, and reinforcement learning (RL) each have independent directories, independent cleanup strategies, independent manifests, and independent execution contexts.

[0048] The "layout signature" mentioned in this invention refers to: performing a hash fingerprint or equivalent encoding on the structure of a round of layout results, which is used to determine whether different rounds of layout are essentially duplicated.

[0049] The "dynamic operational experience base" mentioned in this invention refers to a set of short-term experience contexts that are continuously accumulated during reinforcement learning sessions based on automatic checks and human feedback.

[0050] The "deterministic schema" described in this invention refers to a structured protocol that imposes strong constraints on the output fields, field types, indexing rules, and prohibition of additional fields of a large model.

[0051] The "stagnation rounds" mentioned in this invention refer to the number of consecutive rounds in which the current best score or the score of the most recent rounds is not exceeded, which is used to trigger stronger exploration.

[0052] The "preference sample" mentioned in this invention refers to a data record containing positive samples, negative samples, or pairwise comparison relationships, which can be used for preference optimization training such as Group Relative Policy Optimization (GRPO) and Kahneman-Tversky Optimization (KTO).

[0053] The "teacher model" mentioned in this invention refers to a large-parameter model with strong general reasoning and assembly generation capabilities.

[0054] The "student model" mentioned in this invention refers to a small-to-medium parameter model that is used for low-cost deployment after learning the capabilities of the teacher model through distillation.

[0055] The "direct download CAT Product link" described in this invention refers to the complete process of creating a temporary assembly in the background, saving it as a downloadable file, closing the temporary document, and restoring the original active document without polluting the user's current active document.

[0056] The "large model" mentioned in this invention includes, but is not limited to, "large language model" and "multimodal large model".

[0057] The "large language model" mentioned in this invention refers to a large-scale language model (LLM), which is a language model with a large number of parameters. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks, including text summarization, translation, and sentiment analysis.

[0058] The “Multimodal Large Models” mentioned in this invention refer to models that combine multimodal information such as text, images, videos, and audio for training, including but not limited to multimodal large language models.

[0059] Please see Figure 1 , Figure 1 This is a flowchart of a method for achieving self-evolution of a three-dimensional assembly intelligent agent through computing power using an intelligent computing cloud platform, as provided by the present invention. Figure 1 As shown, the method includes: Step S1: Schedule the computing power in the intelligent computing cloud platform to generate configuration data for each business module based on the 3D files uploaded by the user for multiple business modules.

[0060] The aforementioned business modules include an assembly module, an inspection module, and a reinforcement learning module. The assembly module is responsible for part reading, geometric summary extraction, rule layout generation, and large model layout generation. The inspection module is responsible for product tree construction, distance estimation, experience assembly, report generation, and score interpretation. The reinforcement learning module is responsible for session state maintenance, exploration strategy selection, stagnation detection, diversity transformation, and optimal solution replay.

[0061] The aforementioned 3D files can be CATPart parts or CATProduct assembly files.

[0062] In this invention, the 3D files uploaded by the assembly module can provide the geometric shape, parametric features, and origin coordinate system information of the smallest independent unit, i.e., the parts, for assembly, or they can provide the hierarchical structure, spatial positional relationship, and assembly constraints between the parts.

[0063] The 3D files uploaded by the inspection module can provide a basis for quality assessment and compliance verification of the layout scheme, as well as a complete component topology relationship for performing collision calculations, etc.

[0064] The 3D files uploaded by the reinforcement learning module can provide a simulated physical environment for the reinforcement learning agent to conduct trial and error and evolve, and provide the iteration termination conditions for reinforcement learning in the current 3D assembly agent self-evolution scenario.

[0065] The aforementioned configuration data serves as the basis for the execution of each business module. It may include status information such as the task load to be executed by the business module, and metadata such as the filename and path of the 3D files uploaded by the business module. In this invention, the configuration data can be stored in the intelligent computing cloud platform in the form of a silent task directory and a manifest file.

[0066] In one embodiment, the intelligent computing cloud platform can receive user-uploaded 3D files and establish a three-module silent mode. Separate task directories, cleanup strategies, and runtime contexts are established for each of the three business modules: automated assembly, intelligent inspection, and reinforcement learning. User-uploaded CATPart or CATProduct files are written to the corresponding module's silent task directory, generating a manifest file to obtain configuration data. Subsequent calls can read the task payload through the silent task identifier `silentTaskId` in the silent task directory, without relying on the active window state of current enterprise-level industrial software such as Computer Aided Three-dimensional Interactive Application (CATIA).

[0067] In one embodiment, the manifest file may include, but is not limited to, fields such as task_id, module, created_at, ttl_hours, files[], and manifest_version. The files[] may contain information such as original_name, stored_name, file_type, size_bytes, relative_path, and checksum.

[0068] For example, enterprise users can upload multiple CATPart files to the assembly module via a webpage. The intelligent computing cloud platform creates a task directory under the storage / assembly directory family corresponding to the assembly module, writes it to a manifest file, and generates a silent task identifier.

[0069] It should be noted that, in the embodiments of this invention, the inputs of different business modules can be set with different priorities when calling different business modules. For the automatic assembly module, the explicitly passed silentTaskId can be read first, followed by the explicitly passed parts, and then an attempt can be made to read the currently opened CATIA parts. For the inspection module, the silentTaskId is read first, followed by the externally passed product, and finally the current product. For the reinforcement learning module, the task corresponding to the silent session is read first, and then it falls back to the front-end cached parts or the current CATIA. Through explicit priority design, service-oriented operation and compatibility with old modes can be guaranteed to coexist.

[0070] The beneficial technical effects achieved by step S1 are as follows: The intelligent computing cloud platform stores assembly, inspection, and reinforcement learning tasks separately in isolation, which can prevent different business processes from contaminating each other. This allows both the automatic layout of 3D assembly and the execution of inspection tasks to be carried out in independent silent directories. This enables 3D assembly, inspection, and reinforcement learning tasks to run independently of the current CATIA window state, significantly improving concurrency and stability, and preventing high-load computing or abnormal crashes in the background from affecting the user's currently editing document.

[0071] Step S2: Based on the configuration data corresponding to the assembly module, obtain the feature information of the parts to be assembled, process the feature information, and generate candidate layout schemes.

[0072] The aforementioned features include, but are not limited to, geometric summary, quality, centroid, bounding size, and naming information.

[0073] In one embodiment, based on the configuration data corresponding to the assembly module, multiple tasks that require the generation of layout schemes for the assembly module can be determined. For each task, the intelligent computing cloud platform can allocate computing power to determine the 3D file input part corresponding to the task from the configuration data and extract feature information such as geometric summary, mass, center of gravity, bounding dimensions, and naming information.

[0074] It should be noted that in this invention, for different 3D assembly scenarios with different requirements, feature information of different quantities or precision can be extracted to meet the requirements of low latency and high precision in 3D assembly.

[0075] The above layout scheme describes the final spatial state of a group of parts or subassemblies after they are reasonably arranged in a given space according to specific engineering constraints (such as anti-interference, safety distance, functional coordination, etc.).

[0076] In one embodiment, the intelligent computing cloud platform can invoke a rule-based scheme generator or a large-model layout generator to process the extracted feature information and generate candidate layout schemes. It should be noted that, in this invention, when generating candidate layout schemes from a large model, deterministic schema constraints can be used on the output fields, or multi-stage suggestive degradation can be employed to make the large model output easier to parse and apply, reducing the failure rate in large-scale part scenarios.

[0077] In one embodiment, candidate layout schemes may include fields such as round_index, anchor_index, pose set placements, layout_signature, parse_stage, runtime_context, is_transformed, and score. Each element in the placements array includes at least part_index, relative_to_index, side assembly direction, gap_mm (in millimeters), and optional rotation information, confidence, and source_tagt.

[0078] For example, after receiving an automatic assembly request carrying a silent task identifier and the parameter `useLLM=true` indicating the need to call a large model, the intelligent computing cloud platform can utilize computing power to read all parts in the background, extract geometric summaries and other feature information from all parts, and then call the deterministic layout generator to output placements. The entire process is independent of whether the user has CATIA open or the user's current interface. This embodiment can be used for nighttime batch calculations, remote assembly via a browser, and concurrent access by multiple users.

[0079] It should be noted that for the part feature data extracted for the same task, the rule scheme generator or the large model layout generator can be called to generate multiple candidate layout schemes. Then, the distributed computing power of the intelligent computing cloud platform can be called to check all candidate layout schemes and perform reinforcement learning in parallel.

[0080] The beneficial technical effects achieved in step S2 are as follows: relying on the elastic computing power of the intelligent computing cloud platform, the entire chain from feature extraction to feature processing and then to candidate solution generation runs in an isolated sandbox in a silent task directory. Essentially, it upgrades the traditional manual trial and error and serial operation mode in single-machine CAD assembly to a parallel exploration, automatic screening, and data-driven industrial intelligent assembly mode. This not only greatly improves the efficiency and quality of assembly design, but also provides a solid data foundation and computing power guarantee for the self-evolution system of 3D assembly intelligent body.

[0081] Step S3: Based on the configuration data corresponding to the inspection module, inspect the candidate layout schemes, obtain the inspection results, and update the state parameters of the current session.

[0082] The above inspection results may include, but are not limited to, total score, sub-item evaluation, risk points, recommendations, and explanatory text.

[0083] In one embodiment, based on the configuration data corresponding to the inspection module, the content to be inspected for candidate layout schemes and the scoring criteria used for the inspection can be determined. The intelligent computing cloud platform can allocate computing power to perform an inspection chain on the candidate layout schemes, including product tree acquisition, part summary collection, distance information estimation, experience base assembly, and large model structured analysis, to obtain the inspection results.

[0084] In one embodiment, after obtaining the inspection results, an inspection report can be generated and stored in the database. The inspection report may include, but is not limited to, fields such as report_id, session_id, task_id, score_1_to_10, risk_items, suggestions, explain, component_summary, distance_summary, and created_at. The automatic inspection results not only serve as a user-readable report but also as a source of reward signals and preference labels in reinforcement learning.

[0085] The aforementioned current session state parameters refer to a set of aggregated metrics that are continuously maintained and updated throughout the lifecycle of a user session, used to describe the iterative optimization process of the layout scheme. State parameters include at least the historical best score (best_score), the round number of the best-scoring scheme (best_round_index), score history (scorehistory), stagnation rounds (stagnation rounds), and layout signature counts (layout signature counts).

[0086] In some embodiments, after each round of checks, the current session's state parameters can be updated based on the check results corresponding to the newly generated layout scheme, determining whether the optimal score has changed and whether the number of stalled rounds has increased. For example, if the scores are the same in recent rounds or the score in the new round does not exceed the historical best, the stall count is increased; or, if the layout signatures of candidate layout schemes are repeated, it is considered structural homogenization, and the count of the occurrence of that layout signature is incremented by 1.

[0087] The beneficial technical effects achieved by step S3 are as follows: By combining the elastic computing power and distributed architecture of the intelligent computing cloud platform, the inspection tasks of multiple candidate layout schemes in the same session can be decomposed into independent sub-tasks and distributed to different computing nodes for parallel execution, thereby improving the inspection efficiency of automatically generated layout schemes, perceiving the optimization process in real time, and providing a reliable basis for converging to a higher quality assembly layout scheme.

[0088] It should be noted that, in this invention, the inspection scoring method can be extended from a rule-based plus model hybrid approach to a pure learning inspector or a multi-model integrated inspector.

[0089] Step S4: Based on the inspection results, the current session's state parameters, and candidate layout schemes, determine the exploration strategy and reward function value, and update the dynamic running experience base.

[0090] The exploration strategies mentioned above can include baseline exploration, focused local exploration, diversified global exploration, and aggressive exploration.

[0091] In one embodiment, because the states of historically generated layout schemes differ during reinforcement learning, different optimization directions or exploration efforts need to be determined to identify better candidate layout schemes in order to find the global optimal solution within limited computing power and time. Therefore, the exploration strategy for the next round can be dynamically selected based on candidate schemes, inspection results, and session state parameters.

[0092] Understandably, the reinforcement learning part of this invention is not simply "letting the model try a few more times," but rather explicitly constructing a self-evolving state machine that, like an engineer, can remember effective experiences, identify stagnation, proactively change strategies, and accumulate training assets through multiple rounds of trials.

[0093] In one embodiment, the current session's state parameters can be used to determine whether the number of stalled rounds is too high, whether the layout signature is duplicated, etc., and to select an appropriate exploration strategy. For example, if the number of stalled rounds is less than a certain threshold, only fine-tuning is needed, so focused exploration can be selected; conversely, if the number of stalled rounds exceeds a certain threshold or there are several rounds of duplicate layout signatures, the system is trapped in severe local optima and requires stronger exploration, so strong exploration can be selected.

[0094] The aforementioned reward function value transforms the inspection results into a quantified scalar signal that the agent can understand and learn. It determines how the policy network parameters are updated during the agent's self-evolution. A higher reward indicates a greater tendency to repeat the previous action pattern; a lower reward indicates a greater tendency to explore other directions. In one embodiment, the reward function for reinforcement learning can be defined based on the score in the inspection results, existing risk penalties, and the structural legality reward of the layout scheme.

[0095] The aforementioned dynamic operational experience base, also known as the dynamic experience library, is a core data structure used to store the history of an agent's interactions with the assembly environment during reinforcement learning. An experience item is an interaction record stored in the dynamic operational experience base. After generating candidate layout schemes, each candidate layout scheme can generate a corresponding experience item, recursively updating the dynamic operational experience base. Experience items that can significantly improve the score or successfully break through stagnation have higher weights, while experience items that appear repeatedly in multiple rounds but offer no benefit have lower weights.

[0096] The beneficial technical effects achieved by step S4 are as follows: The computing power based on the intelligent computing cloud platform can ensure the high-speed operation of strategy decision-making and experience base maintenance. By exploring strategy decision-making and reward function calculation and dynamically updating the experience base, it can perceive the current situation, decide the next direction, remember historical experience and continue to evolve. It can proactively improve the assembly quality in multiple rounds, rather than staying at the level of one-time suggestions, and reduce the computing power cost required for the convergence of higher quality assembly layout schemes.

[0097] Step S5: Based on the exploration strategy and reward function value, generate new candidate layout schemes, obtain the inspection results and state parameters corresponding to the new candidate layouts, and construct preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met.

[0098] In one embodiment, after determining the exploration strategy and reward function value, a potentially better candidate layout scheme can be generated based on the exploration strategy and reward function value to further determine whether the exploration strategy is reasonable, thereby enabling the self-evolution of 3D assembly.

[0099] In one embodiment, the 3D file uploaded to the reinforcement learning module may contain the iteration termination condition of reinforcement learning, such as the maximum number of iterations, the number of candidate layout schemes generated, etc. If the iteration termination condition is met, it can be assumed that all candidate layout schemes generated in the past contain the optimal layout scheme, and there is no need to continue learning.

[0100] The aforementioned preference samples may include, but are not limited to, fields such as sample_id, session_id, round_index, source_kind, decision context prompt, layout response, score, category label (e.g., positive or negative), chosen, rejected, and created_at. For manually verified samples, it may also include manual_score, manual_summary, manual_experience, and manual_layout.

[0101] In one embodiment, preference samples can be generated based on automatic inspection scores and human confirmation results. First, rounds with the highest scores in automatic inspection exceeding a target threshold can be marked as positive samples, and the rest as negative samples. Furthermore, candidate layout schemes that have been manually confirmed can be marked as positive. For candidate layout schemes marked as positive in both manual and automatic inspections, the larger value between the automatic score and the human score is used as the valid score. Preference samples can be constructed for all candidate layout schemes for model training. These preference samples can be output in KTO single-label format, or GRPO samples in chosen / rejected form can be constructed in batches.

[0102] In one embodiment, for multiple rounds of candidate layout schemes within the same run batch The system can automatically check the score, manually correct the score, and filter the chosen and rejected options based on the structural legality of each candidate layout scheme. This allows us to obtain the GRPO sample for each running batch, as shown in formula (1).

[0103] (1) In formula (1), For the assembly task context, This indicates the preferred layout or explanation within this batch. This indicates a poor layout or interpretation within that batch. Through this sample format, the student model can learn which layout strategy it prefers under the same task conditions.

[0104] In one embodiment, for a single round of samples, a KTO sample can be constructed as shown in the following formula (2).

[0105] (2) In formula (2), The classification of a sample as negative or positive can be determined based on the target score threshold, the relative best score, and the results of manual verification. The KTO sample type is suitable for use when there is only a single result but with a clear good or bad label.

[0106] In this invention, although all layout schemes are automatically generated, the contribution of different layout schemes to the policy network is quite different. Therefore, after constructing the preference samples, the credibility of each preference sample can be estimated. The credibility weight assigned to each preference sample is calculated by the following formula (3): (3) In formula (3), Indicates the inspection score. This indicates a penalty for parsing uncertainty or missing fields. This indicates whether it has been manually confirmed. , , The weighting coefficients can be customized according to actual needs. The higher the confidence level of the preferred sample, the greater its weight in distillation training.

[0107] In one embodiment, a mechanism for storing preference samples and datasets on disk can be established. Training samples and preference samples are generated based on automatic inspection results and manual confirmation results, and can be exported as datasets in formats such as GRPO and KTO. Simultaneously, they are automatically stored on disk according to task directories and session directories, providing continuously accumulating data assets for subsequent offline training.

[0108] For example, a user can upload a directory containing subfolders in the reinforcement learning module, with each subfolder acting as an independent silent task. The intelligent computing cloud platform can sequentially execute the upload, initiate an automatic session, generate a layout, perform automatic checks, update the memory, synchronize preferred samples, persist the dataset, and then proceed to the next task. If a task achieves identical scores and layout signatures in three consecutive rounds, the optimization stagnation event (stagnation_detected) is triggered. This switches the exploration strategy for generating the next round's layout scheme from focused to aggressive, and applies rotation and spacing scaling transformations. Ultimately, a higher-scoring scheme can be automatically found, and the chosen / rejected samples are stored as GRPO data.

[0109] For example, in manual mode, each round of automated checks can be paused to allow engineers to submit human scores, their experience, and optional manual layouts. The intelligent computing cloud platform can then use the larger of the automated and human scores as the final valid score and mark that round of samples as manually verified positive. Compared to purely automated samples, expert experience can be directly injected into the training assets, improving the professionalism of subsequent distillation models.

[0110] The beneficial technical effects achieved in step S5 are as follows: By constructing preference samples, multiple candidate layout schemes generated in parallel by computing power on the intelligent computing cloud platform can be used for model training through group comparative learning, providing a clearer relative advantage signal for subsequent model training, which helps to improve the efficiency of model training and reduce computing power consumption.

[0111] It should be noted that in this invention, files such as GRPO format preference sample grpo_latest.json, KTO format single sample label data kto_latest.json, and preference dataset latest_dataset.json can be automatically written to the silent task directory. Furthermore, session-specific versions can be saved in the session_id subdirectory to ensure that a retrainable data asset is formed after each batch processing task is completed.

[0112] Step S6: Based on the candidate layout scheme, reward function value, dynamic running experience library and preference sample, the student model is distilled to obtain the target assembly model.

[0113] In one embodiment, a distillation compression link can be established from the teacher model to the student model. The intelligent computing cloud platform uses a large model or a large policy model as the teacher model and a student model with 32B or smaller parameters as the carrier. Through structural distillation, policy distillation, preference distillation, and reward distillation, the assembly layout generation capability, inspection and interpretation capability, and stagnation exploration capability are transferred to the student model, thereby training a target assembly model that can support low-cost deployment in cloud centers, edge nodes, campus all-in-one machines, or local private network environments.

[0114] For example, an enterprise uses a large cloud-based model as the teacher, collecting GRPO / KTO data, inspection reports, and dynamic experience summaries generated from several weeks of silent reinforcement learning sessions to train a 32B student model. After training, the student model retains most of its structured output capabilities and stagnation avoidance capabilities for common assembly tasks, while inference latency and deployment costs are significantly reduced. This makes it suitable for inclusive intelligent computing centers to provide deployable assembly intelligence services to multiple customers.

[0115] In some possible embodiments, after obtaining the target assembly model, continuous distillation can also be achieved based on the intelligent computing cloud platform. That is, after the target assembly model is deployed to the production environment, new layout schemes and user feedback in actual use are continuously collected. This data is fed back to the dynamic experience base of the intelligent computing cloud platform and aggregated with historical data. The teacher model is incrementally trained regularly (e.g., weekly) using the new data, and then the updated teacher model is distilled to generate a new generation of target assembly model.

[0116] The beneficial technical effects achieved by step S6 are as follows: By performing distillation training in the intelligent computing cloud platform, a target assembly model that can be deployed in a lightweight manner is obtained. This ensures that the 3D assembly has a high success rate, strong structural legality and better inspection score, while significantly reducing computing power costs, adapting to different computing power conditions and enhancing the feasibility of industrialization.

[0117] In this invention, step S1 involves scheduling computing power in an intelligent computing cloud platform to generate configuration data for each business module based on the 3D files uploaded by the user for multiple business modules. These business modules include an assembly module, an inspection module, and a reinforcement learning module. Step S2 involves obtaining feature information of the parts to be assembled based on the configuration data for the assembly module, processing the feature information, and generating candidate layout schemes. Step S3 involves inspecting the candidate layout schemes based on the configuration data for the inspection module, obtaining inspection results, and updating the state parameters of the current session. Step S4 involves determining the exploration strategy and reward function value based on the inspection results, the state parameters of the current session, and the candidate layout schemes, and updating the dynamic running experience base. Step S5 involves generating new candidate layout schemes based on the exploration strategy and reward function value, obtaining the inspection results and state parameters corresponding to the new candidate layouts, and constructing preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met. Step S6 involves distilling the student model based on the candidate layout schemes, reward function values, the dynamic running experience base, and preference samples to obtain the target assembly model. In this way, the present invention can rely on the computing power scheduling and unified service capabilities of the intelligent computing cloud platform to automatically generate a layout scheme for three-dimensional assembly, forming a replicable technical path for industrial enterprises. The layout scheme is then intelligently inspected to obtain inspection results. Reinforcement learning is used to achieve online self-evolution of the assembly strategy. Furthermore, by constructing preference samples, the assembly capabilities of the large model are distilled into small and medium-sized models that are more suitable for industrial implementation. This significantly reduces the model inference cost and the engineering deployment threshold while ensuring the success rate of automatic assembly and the interpretability of inspection.

[0118] The intelligent computing cloud platform provided by this invention enables the self-evolution of 3D assembly intelligent agents through computing power. It can be applied to the 3D assembly design and digital verification process in industries such as automotive, aerospace, equipment manufacturing, electronic equipment, complex tooling, and engineering machinery. Specifically, it can include, but is not limited to, automatic assembly pre-layout and rapid review for design departments; assembly risk inspection, guide gap verification, and batch processing verification for process departments; offline RL training, preference sample management, and small model distillation deployment for model governance departments; assembly intelligent service hosting, student model distribution, and cloud-edge collaborative inference for computing power centers; and unified industrial intelligent agent interface capability output for enterprise digital platforms.

[0119] From an industry implementation perspective, this invention can be deployed as an internal enterprise system or as a standard cloud service. Because it includes data accumulation, distillation and compression, and multi-level computing power adaptation mechanisms, it can be deeply integrated with intelligent computing cloud platform services: using stronger teacher models and larger training resources on the central side, and using distilled student models on the edge side to handle high-frequency production tasks, thereby forming a sustainable and commercially viable product system.

[0120] In one embodiment, step S4 includes: Step S4.1: Determine the exploration strategy based on the current session's state parameters.

[0121] The aforementioned state parameters include the current best score, historical score sequence, stagnation count, layout signature set, and dynamic experience base summary.

[0122] In some embodiments, rule-based gating can be used to set a threshold for each parameter in the state parameters. The corresponding exploration strategy is determined based on the relationship between the values ​​of the current session's state parameters and the threshold. Alternatively, the state parameters of the current session can be encoded, and the selection probability corresponding to different exploration strategies can be calculated based on the encoding results. The exploration strategy is then determined based on the probability magnitude. A stagnation index can also be calculated based on the historical score sequence and the current best score in the state parameters. If the stagnation index exceeds a corresponding threshold, a high-intensity exploration strategy is triggered. Furthermore, if the layout signature of a currently generated candidate layout scheme appears multiple times in the layout signature set, the exploration strategy can be determined to perform transformations such as mirroring, rotation, spacing scaling, anchor point offsetting, main axis switching, Z-shaped perturbation, and local expansion on the candidate layout scheme to achieve proactive deduplication.

[0123] Step S4.2: Determine the reward function value based on the score and number of risk penalty items in the inspection results, as well as the structural legality index of the candidate layout scheme.

[0124] In some embodiments, the reward function value is calculated as shown in formula (4): (4) In formula (4), To automatically check the output score of the candidate layout schemes generated in round t, it can be obtained by mapping from a score of 1 to 10. The number of risk penalty items can be composed of the number of high-risk items, the number of intervention items, and the number of unresolved constraints. This is a structural validity indicator for candidate layout schemes. It can consist of indicators such as successful JSON parsing, field completeness, and whether the scheme can be successfully applied.

[0125] In this invention, the model optimization objective no longer focuses solely on "whether it resembles an answer," but is directly linked to industrial feasibility.

[0126] Step S4.3: Based on the inspection results, determine the set of experience items for this layout, and perform fusion calculation with the historical dynamic operation experience library to obtain the updated dynamic operation experience library.

[0127] In one embodiment, the set of empirical terms can be summarized based on all candidate layout schemes generated in round t. The dynamic experience base can be recursively updated according to the following formula (5).

[0128] (5) In formula (5), For experience fusion function, The weights are calculated based on freshness, risk level, and magnitude of improvement. Preferably, experience items that significantly improve scores or successfully break through stagnation have higher weights, while experience items that appear repeatedly in multiple rounds but offer no benefit have lower weights.

[0129] The beneficial technical effects achieved by steps S4.1 to S4.3 are: while significantly reducing computing power costs, the convergence process of the globally optimal assembly scheme is significantly accelerated.

[0130] In one embodiment, the exploration strategy can be determined in multiple ways, then step S4.1 includes at least one of the following: Step S4.1.1: Calculate the selection probability of each candidate exploration strategy based on the current session state parameters, and determine the candidate exploration strategy with the highest selection probability as the target exploration strategy.

[0131] The candidate exploration strategies mentioned above include baseline exploration, local exploration, global exploration, and strong exploration.

[0132] In one embodiment, the state parameters of the current session can be represented by a vector as shown in formula (6).

[0133] (6) In formula (6), This indicates the current best score. This represents a summary of the historical score sequence. Indicates a stalled count. This represents the set of most recently laid-out signatures. This represents a summary of the dynamic experience base. The intelligent computing cloud platform can be based on state parameters. Choose an exploration strategy The probability of choosing an exploration strategy can be calculated using the following formula (7): (7) In formula (7), For the output of the state encoder, , These are the parameters to be learned.

[0134] The beneficial technical effects achieved by step S4.1.1 are as follows: Since different exploration strategies consume computing power in different ways, the computing power allocation ratio can be dynamically adjusted according to the current state through probabilistic strategy selection on the intelligent computing cloud platform, making the computing power allocation more refined and improving the computing power utilization rate.

[0135] Step S4.1.2: Based on the current best score and the historical score sequence, calculate the stagnation index. If the stagnation index exceeds the first threshold, determine the target exploration strategy as strong exploration.

[0136] The aforementioned first threshold can be a value that can be customized according to actual needs, and this invention does not limit it.

[0137] In one embodiment, the current best score is denoted as Based on the historical score sequence, the inspection score for the most recent k rounds is... Then the stagnation index can be calculated using the following formula (8). .

[0138] (8) In formula (8), This indicates the recent signature duplication rate. , , For weights. When When the threshold is exceeded, a high-intensity exploration can be triggered.

[0139] In this invention, when the stagnation index exceeds the first threshold, it indicates that the current exploration strategy can no longer produce substantial progress. Continuing to invest computing power in the original direction will only produce a large number of repetitive or degraded candidate solutions. Therefore, the target exploration strategy can be determined to be strong exploration, actively moving away from the current convergence region to explore the solution space that has never been touched before.

[0140] The beneficial technical effects achieved by step S4.1.2 are: it can avoid the computing power of the intelligent computing cloud platform from being wasted in local optimal areas, avoid ineffective computing power consumption, ensure that computing power is always used in the most valuable search direction, and significantly improve the efficiency of finding the global optimal solution and the convergence quality.

[0141] Step S4.1.3: Based on the number of times the layout signature corresponding to the candidate layout scheme appears in the layout signature set, if the number exceeds the second threshold, determine the target exploration strategy as transforming the candidate layout scheme.

[0142] The aforementioned second threshold can be a value that can be customized according to actual needs, and this invention does not limit it.

[0143] In one embodiment, the candidate layout scheme can be composed of a set of placements, and each placement can be encoded to obtain the encoding result shown in the following formula (9).

[0144] (9) In formula (9), For part_index, For relative_to_index, In the side direction, The result is the discretized gap_mm. Then, the entire layout can be sorted and concatenated before hashing to obtain the layout signature as shown in formula (10): (10) In one embodiment, after obtaining the layout signature corresponding to the candidate layout scheme according to the above formula (10), the number of times the layout signature has appeared in the layout signature set can be determined. If the number exceeds the second threshold, it indicates that the candidate layout scheme is repeated. At this time, the intelligent computing cloud platform can perform transformations such as mirroring, rotation, spacing scaling, anchor point offsetting, main axis switching, Z-shaped perturbation, and local expansion on the candidate layout scheme at the rule level to actively break the homogeneity. Thus, it no longer passively relies on the randomness of the large model, but strengthens the effectiveness of exploration through structural transformation. Specifically, the intelligent computing cloud platform can call computing power to apply transformation factors to the candidate layout scheme. The transformation of the candidate layout scheme can be represented by the following formula (11): (11) In formula (11), It can represent mirror axis, rotation angle, spacing scaling ratio, anchor point offset, etc. By introducing... It can generate new candidate solutions while maintaining some structural rationality. The beneficial technical effect achieved by step S4.1.3 is as follows: When the number of times a certain signature appears in the set exceeds the second threshold, it indicates that the system is repeatedly generating the same type of layout and has fallen into an invalid loop. At this time, the target exploration strategy is switched to transform the candidate layout scheme, forcibly applying perturbation to the current layout scheme, ensuring the diversity coverage of the solution space, and deduplication and mutation can be completed at the front end of the computing power consumption link, avoiding the waste of back-end verification resources.

[0145] In one embodiment, after step S4.2, the method further includes: Step S4.4: In manual mode, obtain the human scores for the candidate layout schemes.

[0146] Step S4.5: Based on the human scoring, the coefficients of the human scoring mapped to the unified reward space and the reward function value are used to determine the reward function value after the human feedback is fused.

[0147] In one embodiment, in manual mode (or human mode), the system can pause after each round of automated checks, waiting for engineers to submit human scores, experience assessments, and optional manual layouts. The human scores are recorded as follows: Combined with the reward function value obtained from automatic inspection and calculation The coefficient that maps human scoring to a uniform reward space The reward function value after artificial feedback fusion can be calculated using the following formula (12).

[0148] (11) The beneficial technical effects achieved by steps S4.1 to S4.3 are: human experience can be written into the dynamic experience base and preference samples, enabling the intelligent computing cloud platform to have enhanced human-machine collaboration capabilities when achieving automatic 3D assembly through computing power.

[0149] In one embodiment, step S6 includes: Step S6.1: Based on candidate layout schemes, reward function values, dynamic running experience base, and preference samples, determine at least two of the distillation losses corresponding to supervised distillation, preference distillation, reward distillation, and context distillation.

[0150] Step S6.2: The distillation loss is weighted and summed to obtain the comprehensive loss. The student model is trained using the comprehensive loss to obtain the target assembly model.

[0151] In one embodiment, the teacher model is: The student model is The teacher model receives the task context x and outputs a summary of the layout, explanation, or strategy. Student model output This invention trains the student model by minimizing the combined loss between the student model output and the teacher model output, thereby reducing the size of student parameters and inference costs while preserving the teacher's capabilities as much as possible. For example, after calculating the distillation losses of four types—supervised distillation, preference distillation, reward distillation, and contextual distillation—the combined loss can be calculated using the following formula (12).

[0152] (12) In formula (12), Indicates supervision of distillation, Indicates preference for distillation. Indicates reward distillation, and This indicates contextual distillation. The weighting coefficients for each type of distillation can be customized according to actual needs, and the comparison in this invention is not limited.

[0153] In this invention, a student model is trained using structured layouts, inspection explanations, experience summaries, preference selections, and exploration trajectories generated by a teacher model. The distillation process can include four types of objectives: supervised distillation, preference distillation, reward distillation, and contextual distillation. This enables the student model to not only learn "to output a certain layout," but also to learn "why it is arranged this way, when to break out of stagnation, and how to correct the strategy from the inspection results," resulting in a target assembly model that can be used for low-cost deployment.

[0154] The beneficial technical effects achieved by steps S6.1 and S6.2 are as follows: Through multi-objective distillation, the capabilities originally dependent on high-cost teacher models can be compressed into smaller student models. This reduces deployment difficulty while maintaining high success rates, strong structural validity, and superior check scores in the 3D assembly domain. This enables the model to be deployed on enterprise edge computing nodes, dedicated inference services, and even offline environments, significantly reducing computing costs.

[0155] In one embodiment, step S6.1 includes at least two of the following: Step S6.1.1: Determine the supervised distillation loss based on the candidate layout scheme.

[0156] In one embodiment, the structured JSON output by the teacher model, i.e. the candidate layout scheme, can be used as the target. The supervised distillation loss can be calculated using the following formula (13) so that the student model can generate a layout result that is consistent with the teacher's style and has complete fields.

[0157] (13) Step S6.1.2: Determine the preference distillation loss based on the preference sample.

[0158] In one embodiment, the preference distillation loss can be calculated using the chosen / rejected paired samples using the following formula (14).

[0159] (14) In formula (14), The scaling factor, the preference distillation loss, can make the student model more biased towards high-scoring layouts.

[0160] Step S6.1.3: Determine the reward distillation loss based on the reward function value.

[0161] In one embodiment, the student model can also be fitted with a reward head or implicit score. Therefore, the reward distillation loss can be calculated using the following formula (15).

[0162] (15) In formula (15), For reward estimates defined by teachers or external checkers, For student reward prediction, calculating the reward distillation loss makes it easier for the student model to predict whether a layout is good or bad during inference.

[0163] Step S6.1.4: Determine the context distillation loss based on the dynamic operation experience base.

[0164] In one embodiment, industrial assembly tasks are highly sensitive to context compression. To enable the student model to retain the ability to "which experiences are important and which risks should be prioritized," this invention can also calculate the loss of context summarization distillation, as shown in the following formula (16): (16) In formula (16), This represents the latent variable representations of the teacher and student models for the key context of the task. This represents a summary of the dynamic experience base. Calculating the context distillation loss makes it easier for student models to inherit the teacher model's judgments on risk patterns and exploration directions.

[0165] In one embodiment, step S2 includes: Step S2.1: Based on the current requirements of the 3D assembly scene, determine the feature information extraction mode of the parts to be assembled.

[0166] Step S2.2: Based on the information range corresponding to the extraction mode, extract the feature information of the parts to be assembled from the configuration data.

[0167] In one embodiment, for low-latency scenarios, the system can use a fast extraction mode to extract lightweight geometric summaries and finite components, distances, and empirical upper limits; for high-precision scenarios, a standard mode can be enabled to obtain richer topology and distance information, thus balancing real-time performance and precision in different application scenarios.

[0168] In one embodiment, the intelligent computing cloud platform can determine the appropriate extraction mode (fast mode or standard mode) based on whether the current 3D assembly scene corresponds to the fastmode flag. Then, based on the information range corresponding to the extraction mode, it determines whether to extract lightweight geometric summaries or standard geometric features. In fast mode, heavy topology extraction can be skipped, retaining only dimensions, mass, centroid, naming, and a few summary features to support large models in quickly understanding part relationships. Conversely, in standard mode, more geometric, distance, or contextual features can be extracted to improve recommendation and inspection accuracy.

[0169] For example, to ensure latency in the demonstration environment, enabling fast mode limits the number of components to 32, the distance to 12, and the experience to 12, and uses center of gravity plus size to estimate distance. While the inspection results are slightly coarser than in standard mode, they still retain the interpretability of key risks, allowing the 200-component demonstration task to be completed within an acceptable timeframe.

[0170] The beneficial technical effects achieved by steps S2.1 and S2.2 are as follows: By using the dual-channel feature information extraction mode and selecting the appropriate extraction mode according to the specific needs of the 3D assembly scenario, the dynamic balance between inference efficiency and accuracy can be ensured, and the rationality of computing power scheduling of the intelligent computing cloud platform during 3D assembly can be improved.

[0171] It should be noted that in this invention, the geometric summary in fast mode can be replaced with a point cloud summary, a voxel summary, or a STEP intermediate representation summary.

[0172] In one embodiment, step S2 includes: Step S2.3: Determine the number of parts to be assembled in the 3D model file included in the configuration data.

[0173] Step S2.4: Determine the prompt template for the initial stage based on the relationship between the quantity and the quantity threshold.

[0174] Step S2.5: Generate candidate layout schemes based on the prompt template and feature information from the initial stage.

[0175] Step S2.6: If the candidate layout scheme is non-compliant, update the candidate layout scheme based on the prompt template and feature information of the next stage.

[0176] In one embodiment, multiple prompt templates for different stages can be preset. These prompt templates are ordered from highest to lowest intensity of feature information analysis: rich, ultra, and basic. Therefore, when the number of parts is greater than or equal to a threshold, ultra (compact) prompts are prioritized, then degenerate to basic; when the number of parts is small, rich prompts are prioritized, then degenerate to basic and ultra. The server records the prompt stage hits for subsequent quality analysis.

[0177] For example, when a task contains more than 120 parts, exceeding the proficiency threshold, the initial prompt template is automatically set to "ultra," retaining only key indices, size information, and relationship fields. Large models are required to strictly return `anchor_index` and `placements`, disallowing annotations and extra fields. If the model returns an invalid response on the first attempt, it reverts to "basic"; if this still fails, it reverts to the rule-based layout scheme.

[0178] The beneficial technical effects achieved by steps S2.3 to S2.6 are as follows: Compared with unconstrained natural language prompts, by selecting a suitable prompt template based on the number of parts, the parsing failure rate can be significantly reduced, and the computing power consumption of the intelligent computing cloud platform can be reduced.

[0179] In one embodiment, step S2 includes: Step S2.1': Input the prompt template and feature information containing the output format into the large model to obtain the candidate layout scheme generated by the large model. The output format indicates that the candidate layout scheme is output in JSON mode and includes at least the anchor index and assembly constraint relationship.

[0180] In one embodiment, the intelligent computing cloud platform can allocate computing power to construct a structured deterministic assembly generation mechanism. For automated assembly tasks, the system extracts features such as geometric summary, mass, center of gravity, bounding dimensions, and naming information from the input parts. When calling the large model to generate layout suggestions, the system uses schema to strictly constrain the output fields, ensuring that the generated candidate layout schemes include at least the anchor index, assembly constraint relationships, and fields such as part_index, relative_to_index, side, and gap_mm in each placement, while simultaneously prohibiting additional fields.

[0181] The beneficial technical effects achieved by step S2.1' are: through deterministic schema, hint degradation and field extension prohibition mechanism, the output layout scheme of large model can be more easily parsed and applied, reducing the failure rate in large-scale part scenarios.

[0182] It should be noted that in this invention, JSON Schema constraints can be replaced with Protobuf, XML, or other structured protocols that can guarantee the integrity and parsability of the fields in the layout output.

[0183] In one embodiment, after step S2, the method further includes: Step S7: Upon receiving the download instruction, create a temporary assembly document and write the candidate layout schemes into the temporary assembly document.

[0184] Step S8: With the temporary assembly document saved to the download directory, close the temporary assembly document and restore the original active document.

[0185] In one embodiment, the generated layout can be used for previewing, applying to CATIA, or directly downloading the CATProduct constructed in the backend. During direct download, the intelligent computing cloud platform can create a temporary assembly document in the background, write the layout result, save it as a file to the download directory, close the temporary document, and restore the user's previous active document; if restoration fails, it will be restored by name as a fallback. This avoids the on-site contamination caused by directly rewriting the user's current active document.

[0186] In one embodiment, the intelligent computing cloud platform can establish a sustainable system operation and maintenance framework. The system uniformly provides log rotation, task TTL recycling, module-level isolated storage, dataset persistence, download recovery, and failure fallback mechanisms, ensuring that the solution is not just a theoretical method, but an industrial-grade product capability that can be operated online for a long time.

[0187] This invention is based on silent task isolation, with deterministic assembly generation and intelligent inspection feedback at its core. It is driven by reinforcement learning self-evolution and preference sample accumulation, and extended by teacher-student distillation and universal deployment, forming a complete closed loop encompassing "execution, evaluation, learning, compression, deployment, and relearning." Compared to traditional assembly automation, this invention not only performs one-time automated assembly but also continuously accumulates experience, optimizes strategies, and outputs scalable small-scale models in actual enterprise operations, demonstrating significant novelty, inventiveness, and practicality.

[0188] It should be noted that, in this invention, the layout signature can be extended from the hash method to graph isomorphic signature, constraint graph signature, or learned layout embedding.

[0189] It should be noted that in this invention, the distillation object can be extended from a single student model to a family of multiple student models, including lightweight students, marginal students, and high-precision students.

[0190] It should be noted that in this invention, the RL session can be extended from a single-task serial process to a multi-task parallel process, as long as the session state isolation and sample affiliation are maintained.

[0191] It should be noted that this invention can be adapted to other 3D CAD platforms, as long as the underlying layer has the ability to read, assemble, write, and export parts.

[0192] Please see Figure 2 , Figure 2 It is a device that enables the self-evolution of three-dimensional assembly intelligent agents through computing power in an intelligent computing cloud platform, such as... Figure 2 As shown, the device 20 for the intelligent computing cloud platform to realize the self-evolution of a three-dimensional assembly intelligent agent through computing power includes: The silent task management module 201 is used to schedule computing power in the intelligent computing cloud platform. Based on the three-dimensional files uploaded by the user for multiple business modules, it generates configuration data corresponding to each business module. The multiple business modules include an assembly module, an inspection module, and a reinforcement learning module. The automatic assembly module 202 is used to obtain the feature information of the parts to be assembled based on the configuration data corresponding to the assembly module, process the feature information, and generate candidate layout schemes. The intelligent inspection module 203 is used to inspect the candidate layout scheme based on the configuration data corresponding to the inspection module, obtain the inspection result, and update the status parameters of the current session. The reinforcement learning module 204 is used to determine the exploration strategy and reward function value based on the inspection results, the state parameters of the current session and the candidate layout scheme, and to update the dynamic running experience base. The sample determination module 205 is used to generate new candidate layout schemes according to the exploration strategy and the reward function value, obtain the inspection results and state parameters corresponding to the new candidate layouts, and construct preference samples according to the inspection results and state parameters corresponding to all candidate layout schemes when the iteration termination condition corresponding to the reinforcement learning module is met. The distillation training module 206 is used to distill the student model based on the candidate layout scheme, the reward function value, the dynamic running experience library, and the preference sample to obtain the target assembly model.

[0193] In one embodiment, reinforcement learning module 204 is used for: The exploration strategy is determined based on the state parameters of the current session, including the current best score, historical score sequence, stagnation count, layout signature set, and dynamic experience base summary. The reward function value is determined based on the score and number of risk penalty items in the inspection results, and the structural legality index of the candidate layout scheme; Based on the inspection results, the set of experience items for this layout is determined, and then fused with the historical dynamic operation experience library to obtain the updated dynamic operation experience library.

[0194] In one embodiment, reinforcement learning module 204 is used for at least one of the following: Based on the state parameters of the current session, the selection probability of each candidate exploration strategy is calculated, and the candidate exploration strategy with the highest selection probability is determined as the target exploration strategy. The candidate exploration strategies include baseline exploration, local exploration, global exploration, and strong exploration. Based on the current best score and the historical score sequence, a stagnation index is calculated. If the stagnation index exceeds a first threshold, the target exploration strategy is determined to be strong exploration. Based on the number of times the layout signature corresponding to the candidate layout scheme appears in the layout signature set, if the number exceeds a second threshold, the target exploration strategy is determined to be to transform the candidate layout scheme.

[0195] In one embodiment, reinforcement learning module 204 is further configured to: In manual mode, obtain human scores for the candidate layout schemes; Based on the human scoring, the coefficients of the human scoring mapped to the unified reward space, and the reward function value, the reward function value after human feedback fusion is determined.

[0196] In one embodiment, the distillation training module 206 is used for: Based on the candidate layout scheme, the reward function value, the dynamic running experience base, and the preference sample, determine at least two distillation losses corresponding to supervised distillation, preference distillation, reward distillation, and context distillation, respectively. The distillation loss is weighted and summed to obtain a comprehensive loss. The student model is then trained using this comprehensive loss to obtain the target assembly model.

[0197] In one embodiment, the distillation training module 206 is used for at least two of the following: Based on the candidate layout schemes, the supervised distillation loss is determined; Based on the preference sample, determine the preference distillation loss; Based on the reward function value, determine the reward distillation loss; Based on the aforementioned dynamic operational experience base, the context distillation loss is determined.

[0198] In one embodiment, the automatic assembly module 202 is used for: Based on the current requirements of the 3D assembly scenario, the feature information extraction mode of the parts to be assembled is determined. Based on the information range corresponding to the extraction mode, the feature information of the part to be assembled is extracted from the configuration data.

[0199] In one embodiment, the automatic assembly module 202 is used for: Determine the number of parts to be assembled within the 3D file contained in the configuration data; Based on the relationship between the quantity and the quantity threshold, determine the prompt template for the initial stage; Based on the prompt template from the initial stage and the feature information, candidate layout schemes are generated; If the candidate layout scheme is non-compliant, the candidate layout scheme is updated based on the prompt template for the next stage and the feature information.

[0200] In one embodiment, the automatic assembly module 202 is used for: The prompt template containing the output format and the feature information are input into the large model to obtain the candidate layout scheme generated by the large model. The output format indicates that the candidate layout scheme is output in JSON mode and includes at least anchor index and assembly constraint relationship.

[0201] In one embodiment, the automatic assembly module 202 is further configured to: Upon receiving a download instruction, a temporary assembly document is created, and the candidate layout scheme is written into the temporary assembly document; If the temporary assembly document is saved to the download directory, close the temporary assembly document and restore the original active document.

[0202] The device for realizing the self-evolution of a three-dimensional assembly intelligent agent through computing power provided by the intelligent computing cloud platform of the present invention is capable of realizing the various processes of the various embodiments of the above-mentioned method for realizing the self-evolution of a three-dimensional assembly intelligent agent through computing power. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0203] It should be noted that the device for realizing the self-evolution of the three-dimensional assembly intelligent body through computing power in the intelligent computing cloud platform of this invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0204] The present invention also provides an electronic device, see below. Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 301, a processor 302, and a program or instructions stored in the memory 301 that run on the memory. When the program or instructions are executed by the processor 302, they can achieve the following: Figure 1The corresponding intelligent computing cloud platform achieves any step in the method embodiment of the self-evolution of the three-dimensional assembly intelligent body through computing power and achieves the same beneficial effect, which will not be elaborated here.

[0205] The processor 302 can be a CPU, ASIC, FPGA or GPU.

[0206] Those skilled in the art will understand that all or part of the steps of the above-described method embodiment for realizing the self-evolution of a three-dimensional assembly intelligent body through computing power on the intelligent computing cloud platform can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0207] The present invention also provides a readable storage medium on which a computer program is stored, and which, when executed by a processor, can perform the above-described functions. Figure 1 The corresponding intelligent computing cloud platform achieves any step in the method embodiment of the self-evolution of the three-dimensional assembly intelligent agent through computing power, and can achieve the same technical effect. To avoid repetition, it will not be described again here. The storage medium mentioned is such as read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0208] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The corresponding intelligent computing cloud platform realizes each process of the method embodiment of the self-evolution of the three-dimensional assembly intelligent body through computing power, and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0209] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0210] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0211] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0212] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for implementing self-evolution of a three-dimensional assembly intelligent agent by computing power through an intelligent computing cloud platform, characterized in that, include: Step S1: Schedule the computing power in the intelligent computing cloud platform, and generate configuration data for each business module based on the three-dimensional files uploaded by the user for multiple business modules. The multiple business modules include an assembly module, an inspection module, and a reinforcement learning module. Step S2: Based on the configuration data corresponding to the assembly module, obtain the feature information of the part to be assembled, process the feature information, and generate a candidate layout scheme. Step S3: Based on the configuration data corresponding to the inspection module, inspect the candidate layout scheme, obtain the inspection result, and update the state parameters of the current session; Step S4: Based on the inspection results, the current session's state parameters, and the candidate layout scheme, determine the exploration strategy and reward function value, and update the dynamic running experience base; Step S5: Generate new candidate layout schemes based on the exploration strategy and the reward function value, obtain the inspection results and state parameters corresponding to the new candidate layouts, and construct preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes, provided that the iteration termination condition corresponding to the reinforcement learning module is met. Step S6: Based on the candidate layout scheme, the reward function value, the dynamic running experience library, and the preference sample, the student model is distilled to obtain the target assembly model.

2. The method as described in claim 1, characterized in that, Step S4 includes: Step S4.1: Determine the exploration strategy based on the state parameters of the current session. The state parameters include the current best score, historical score sequence, stagnation count, layout signature set, and dynamic experience base summary. Step S4.2: Determine the reward function value based on the score and number of risk penalty items in the inspection results, and the structural legality index of the candidate layout scheme; Step S4.3: Based on the inspection results, determine the set of experience items for this layout, and perform fusion calculation with the historical dynamic operation experience library to obtain the updated dynamic operation experience library.

3. The method as described in claim 2, characterized in that, Step S4.1 includes at least one of the following: Step S4.1.1: Calculate the selection probability of each candidate exploration strategy based on the state parameters of the current session, and determine the candidate exploration strategy with the highest selection probability as the target exploration strategy. The candidate exploration strategies include baseline exploration, local exploration, global exploration, and strong exploration. Step S4.1.2: Based on the current best score and the historical score sequence, calculate the stagnation index. If the stagnation index exceeds a first threshold, determine the target exploration strategy as strong exploration. Step S4.1.3: Based on the number of times the layout signature corresponding to the candidate layout scheme appears in the layout signature set, if the number of times exceeds the second threshold, determine the target exploration strategy as transforming the candidate layout scheme.

4. The method as described in claim 2, characterized in that, Following step S4.2, the following is also included: Step S4.4: In manual mode, obtain the human scores for the candidate layout schemes. Step S4.5: Based on the human scoring, the coefficients of the human scoring mapped to the unified reward space, and the reward function value, determine the reward function value after human feedback fusion.

5. The method as described in claim 1, characterized in that, Step S6 includes: Step S6.1: Based on the candidate layout scheme, the reward function value, the dynamic running experience base, and the preference sample, determine at least two distillation losses corresponding to supervised distillation, preference distillation, reward distillation, and context distillation respectively; Step S6.2: The distillation loss is weighted and summed to obtain the comprehensive loss. The student model is trained using the comprehensive loss to obtain the target assembly model.

6. The method as described in claim 5, characterized in that, Step S6.1 includes at least two of the following: Step S6.1.1: Based on the candidate layout scheme, determine the supervised distillation loss; Step S6.1.2: Determine the preference distillation loss based on the preference sample; Step S6.1.3: Determine the reward distillation loss based on the reward function value; Step S6.1.4: Determine the context distillation loss based on the dynamic operation experience base.

7. The method as described in claim 1, characterized in that, Step S2 includes: Step S2.1: Based on the current requirements of the 3D assembly scene, determine the feature information extraction mode of the parts to be assembled; Step S2.2: Based on the information range corresponding to the extraction mode, extract the feature information of the part to be assembled from the configuration data.

8. The method as described in claim 7, characterized in that, Step S2 includes: Step S2.3: Determine the number of parts to be assembled in the 3D file contained in the configuration data; Step S2.4: Determine the prompt template for the initial stage based on the relationship between the quantity and the quantity threshold; Step S2.5: Generate candidate layout schemes based on the prompt template from the initial stage and the feature information; Step S2.6: If the candidate layout scheme is non-compliant, update the candidate layout scheme based on the prompt template for the next stage and the feature information.

9. The method according to any one of claims 1 to 8, characterized in that, Step S2 includes: Step S2.1': Input the prompt template containing the output format and the feature information into the large model to obtain the candidate layout scheme generated by the large model. The output format indicates that the candidate layout scheme is output in JSON mode and includes at least anchor index and assembly constraint relationship.

10. The method according to any one of claims 1 to 8, characterized in that, Following step S2, the method further includes: Step S7: Upon receiving a download instruction, create a temporary assembly document and write the candidate layout scheme into the temporary assembly document; Step S8: If the temporary assembly document is saved to the download directory, close the temporary assembly document and restore the original active document.

11. A device for realizing the self-evolution of a three-dimensional assembly intelligent agent through computing power using an intelligent computing cloud platform, characterized in that, include: The silent task management module is used to schedule computing power in the intelligent computing cloud platform. Based on the three-dimensional files uploaded by the user for multiple business modules, it generates configuration data corresponding to each business module. The multiple business modules include an assembly module, an inspection module, and a reinforcement learning module. An automatic assembly module is used to obtain feature information of the parts to be assembled based on the configuration data corresponding to the assembly module, process the feature information, and generate candidate layout schemes. The intelligent inspection module is used to inspect the candidate layout scheme based on the configuration data corresponding to the inspection module, obtain the inspection results, and update the status parameters of the current session. The reinforcement learning module is used to determine the exploration strategy and reward function value based on the inspection results, the state parameters of the current session, and the candidate layout scheme, and to update the dynamic running experience base. The sample determination module is used to generate new candidate layout schemes based on the exploration strategy and the reward function value, obtain the inspection results and state parameters corresponding to the new candidate layouts, and construct preference samples based on the inspection results and state parameters corresponding to all candidate layout schemes when the iteration termination condition corresponding to the reinforcement learning module is met. The distillation training module is used to distill the student model based on the candidate layout scheme, the reward function value, the dynamic running experience library, and the preference sample to obtain the target assembly model.

12. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for achieving self-evolution of a three-dimensional assembled intelligent agent through computing power by the intelligent computing cloud platform as described in any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for achieving self-evolution of a three-dimensional assembled intelligent agent through computing power using the intelligent computing cloud platform as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of the method for achieving self-evolution of a three-dimensional assembly intelligent agent through computing power using the intelligent computing cloud platform as described in any one of claims 1 to 10.