Construction method of industrial large model and storage medium
By building lightweight model tools and knowledge organization structures in the industrial field, fine-tuning multimodal large models, and forming an adaptive routing decision system, the adaptability and real-time performance issues of large models in industrial scenarios are solved, and efficient and reliable industrial applications are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU YIWEI INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-10
AI Technical Summary
Existing large models suffer from low scenario adaptability and high latency in multimodal real-time inference when applied in the industrial field, which limits their application value in core industrial production processes.
By acquiring multimodal fusion data from the industrial field, we classify scenarios, build lightweight model tools and knowledge organization structures, fine-tune and distill large multimodal models and language models, form a scenario model library and a routing central command, achieve adaptive routing decisions, and generate target answers.
It improves the adaptability, real-time performance, and reliability of large industrial models, enabling them to quickly respond to the needs of specific industrial scenarios, reduce inference latency, and enhance the professionalism of the models and the reliability of the output.
Smart Images

Figure CN121835889A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of computer applications and industrial knowledge question-and-answer technology, and in particular to a method for constructing and storing large industrial models. Background Technology
[0002] With the rapid development of artificial intelligence technology, foundational models, represented by Large Language Models (LLM) and multimodal large models, have demonstrated powerful general cognitive and content generation capabilities. Applying large model technology to the industrial field enables intelligent production optimization, equipment operation and maintenance, quality inspection, safety monitoring, and scheduling decisions. However, the in-depth application and implementation of large models in the industrial field faces severe and complex challenges, and existing technologies have significant shortcomings, mainly in the following aspects: 1. Industrial scenarios are highly segmented and fragmented, resulting in low adaptability of general-purpose models: Industrial scenarios cover the entire lifecycle, including R&D design, manufacturing, supply chain management, and after-sales service. Furthermore, the professional knowledge, process specifications, and data patterns vary significantly across different industries (such as semiconductors, automotive, chemicals, and textiles). Existing single-modal large models lack deep knowledge injection and task adaptation capabilities for specific industrial scenarios. This leads to a severe disconnect between output results and industry terminology, process standards, and safety procedures, resulting in poor practicality.
[0003] 2. Large Multimodal Models Experience Slow Inference Speed and High Latency: Industrial environments are inherently multimodal, encompassing sensor time-series signals, machine vision images, video, 3D point clouds, structured process parameters, and unstructured technical documents. Existing large multimodal models are typically bulky and computationally complex, resulting in slow inference speed and high latency when processing high-resolution images, long-series time-series data, or performing complex chained inference. This makes them difficult to integrate into industrial control loops with high real-time requirements, limiting their application value in core industrial production processes. Summary of the Invention
[0004] This disclosure provides a method for constructing an industrial large model and a storage medium to address the problems of existing large model technologies, such as low scene adaptability and large multimodal real-time inference latency, which hinder their large-scale and in-depth implementation in the industrial field.
[0005] Based on the above problems, in a first aspect, the industrial large-scale model construction method provided in this disclosure embodiment, applied to the field of industrial knowledge question answering, includes: Acquire multimodal fusion data from the industrial sector; classify the industrial sector into subdivided industrial scenarios; Based on the multimodal fusion data, determine the multimodal dataset for each industrial sub-scenarios; and based on the multimodal dataset, construct the knowledge organization structure for each industrial sub-scenarios. For each specific industrial scenario, a lightweight model tool is trained based on the multimodal dataset to obtain a scenario model tool; and a scenario model library is constructed based on the scenario model tool. Based on the knowledge organization structure and the scene model library, a database is constructed, which is used to store the input data of the scene model tools; Based on the multimodal dataset, fine-tune the distillation of the large multimodal model to obtain a backup inference body; For each industrial sub-scenarios, a medium-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a scenario expert agent for each industrial sub-scenarios. Based on the text modalities in the multimodal dataset, a small language model is fine-tuned to obtain a routing central command entity. The routing central command entity is used to identify received user instructions and determine the confidence set of the target industrial sub-scenarios. The confidence set of the target industrial sub-scenarios is used to select target models from the operational model library using an adaptive routing decision algorithm, and to generate target answers using the target models. The operational model library includes: a scenario model library, a backup inference entity, and a scenario expert intelligent agent.
[0006] In a second aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, performs the steps of the method for constructing an industrial large model as described in the first aspect or any possible embodiment in conjunction with the first aspect.
[0007] The beneficial effects of the embodiments disclosed herein include: This disclosure provides a method and storage medium for constructing a large industrial model, comprising: acquiring multimodal fusion data in the industrial field; classifying the industrial field into sub-scenarios to obtain sub-scenarios; determining the multimodal dataset for each sub-scenarios based on the multimodal fusion data; constructing a knowledge organization structure for each sub-scenarios based on the multimodal dataset; training a lightweight model tool for each sub-scenarios based on the multimodal dataset to obtain a scenario model tool; constructing a scenario model library based on the scenario model tool; constructing a database based on the knowledge organization structure and the scenario model library, the database being used to store the input data of the scenario model tool; and fine-tuning the distillation multimodal large model based on the multimodal dataset to obtain... The system includes: a backup inference agent; for each industrial sub-scenarios, a medium-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a scenario expert agent for each industrial sub-scenarios; a small-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a routing central command agent; the routing central command agent is used to identify received user instructions and determine the confidence set of the target industrial sub-scenarios; the confidence set of the target industrial sub-scenarios is used to select target models from the operational model library based on the target industrial sub-scenarios confidence set using an adaptive routing decision algorithm, and then using the target models to generate the target answer; the operational model library includes: a scenario model library, a backup inference agent, and a scenario expert agent. The industrial large-scale model construction method provided in this disclosure has significant technical advantages by constructing a multi-entity collaborative and hierarchical evolution industrial large-scale model. Specifically, it is reflected in the high-efficiency collaboration and continuous optimization capabilities across the entire chain, including: 1. A small routing central command unit carried by the industrial large-scale model can realize the rapid intent recognition and accurate scene routing of user commands, laying a solid foundation for high real-time response throughout the entire process; 2. A medium-sized scene expert intelligent agent and a lightweight scene model tool form a high-efficiency collaborative link, outputting professional and accurate solutions for various industrial sub-scenarios, greatly improving the adaptability and accuracy of task processing; 3. When facing complex and difficult tasks, the backup inference agent can conduct in-depth analysis and reasoning, and perform multiple verifications in combination with the knowledge organization structure, effectively suppressing the large model illusion problem and ensuring the reliability of the output target answer. Attached Figure Description
[0008] Figure 1 A flowchart illustrating the method for constructing a large industrial model as provided in this embodiment of the disclosure; Figure 2 A flowchart for generating the target answer provided in this embodiment of the disclosure; Figure 3 A flowchart of a tool for determining a target scene model provided in an embodiment of this disclosure. Detailed Implementation
[0009] This disclosure provides a method for constructing a large-scale industrial model and a storage medium. Preferred embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit this disclosure. Furthermore, the embodiments and features described in this application can be combined with each other unless otherwise specified.
[0010] This disclosure provides a method for constructing a large-scale industrial model, applicable to the field of industrial knowledge question answering, such as... Figure 1 As shown, it includes: S101. Acquire multimodal fusion data in the industrial field; classify industrial scenarios to obtain subdivided industrial scenarios; S102. Based on the multimodal fusion data, determine the multimodal dataset for each industrial sub-scenarios; and based on the multimodal dataset, construct the knowledge organization structure for each industrial sub-scenarios. S103. For each type of industrial sub-scenario, train a lightweight model tool based on a multimodal dataset to obtain a scenario model tool; and build a scenario model library based on the scenario model tool. S104. Based on the knowledge organization structure and scenario model library, construct a database to store the input data of the scenario model tool. S105. Based on the multimodal dataset, fine-tune the distillation of the large multimodal model to obtain a backup inference body; S106. For each type of industrial sub-scenario, fine-tune the medium-sized language model based on the text modalities in the multimodal dataset to obtain the scene expert agent for each type of industrial sub-scenario. S107. Based on the text modalities in the multimodal dataset, fine-tune a small language model to obtain a routing central command entity; the routing central command entity is used to identify the received user instructions and determine the target industrial sub-scenario confidence set; the target industrial sub-scenario confidence set is used to select target models from the operational model library based on the target industrial sub-scenario confidence set using an adaptive routing decision algorithm, and then use the target models to generate target answers; wherein, the operational model library includes: a scenario model library, a backup inference entity, and a scenario expert intelligent entity.
[0011] In this embodiment of the disclosure, with the rapid development of artificial intelligence technology, basic models represented by Large Language Models (LLM) and multimodal large models, with their powerful general cognition and content generation capabilities, have provided important technical support for the intelligentization of production optimization, equipment operation and maintenance, quality inspection, safety monitoring, and scheduling decisions in the industrial field. However, their in-depth application and implementation in industrial scenarios still face severe and complex challenges, and existing technologies have significant shortcomings: On the one hand, industrial scenarios cover the entire life cycle, including R&D design, production and manufacturing, supply chain management, and after-sales service, and the professional knowledge, process specifications, and data patterns of different industries (such as semiconductors, automobiles, chemicals, and textiles) vary greatly, exhibiting highly segmented and fragmented characteristics. Existing single... Large-scale multimodal models lack the ability to inject deep knowledge into specific industrial scenarios and adapt to tasks, resulting in a serious disconnect between output results and industry terminology, process standards, and safety procedures, significantly reducing their practicality. On the other hand, the industrial environment is inherently a multimodal scenario, requiring the processing of various types of data such as sensor time-series signals, machine vision images, videos, 3D point clouds, structured process parameters, and unstructured technical documents. However, existing large-scale multimodal models, due to their large structure and computational complexity, generally suffer from slow inference speed and high latency when processing high-resolution images, long-series time-series data, or performing complex chain inference, making it difficult to integrate them into industrial control loops with high real-time requirements, thus limiting their application value in core industrial production processes.
[0012] In this embodiment, a basic framework is built by acquiring multimodal fusion data and scene classification, constructing a multimodal dataset and knowledge organization structure for subdivided industrial scenarios. Subsequently, lightweight model tools are trained, and the distillation of the large multimodal model and language model are fine-tuned, forming a multi-dimensional model system including scene model tools, backup inference agents, scene expert intelligent agents, and a routing central command agent. Data support is achieved through database construction, and user command recognition, model adaptive matching, and answer generation are completed with the help of the routing central command agent. The entire process incorporates a combat-style architecture design, improving the adaptability, real-time performance, and application reliability of the large industrial model. For example, as shown in Table 1, Table 1 is a comparison table of the large industrial model architecture and the combat-style architecture:
[0013] Table 1 Acquire multimodal fusion data from the industrial sector. This data can include sensor time-series signals, machine vision images, videos, 3D point clouds, structured process parameters, unstructured technical documents, and other data covering the entire industrial landscape. Classify the industrial sector into sub-scenarios. Each sub-scenarios represents the smallest segment that professionals in the field can divide into within a specific stage of R&D, production control, quality inspection, equipment maintenance, procurement and supply, sales and after-sales service, or business management. Examples include equipment anomaly detection and equipment health assessment scenarios within equipment maintenance (analogous to dividing different operational zones like water, land, and air, clarifying the operational needs of each zone). Based on the multimodal fusion data, determine the multimodal dataset for each sub-scenarios. According to the boundaries and requirements of each sub-scenarios, filter, clean, label, and integrate the multimodal fusion data to determine a dedicated multimodal dataset for each sub-scenarios, ensuring a high degree of matching between the multimodal dataset and the core tasks of each sub-scenarios. By combining professional knowledge, process specifications, safety regulations, and technological standards from various industrial sub-scenarios, we identify knowledge relationships and construct a knowledge organization structure for each scenario (such as a knowledge base or knowledge graph, analogous to organizing historical combat reviews and drawing combat maps). This knowledge organization structure enables the systematic accumulation of industrial professional knowledge, providing knowledge constraints for subsequent model training and inference, and preventing model output from deviating from industry standards. For the core tasks of each industrial sub-scenarios (such as quality inspection on a production line or maintenance diagnosis of a piece of equipment), we train lightweight model tools (such as simplified CNNs or Transformer variants) based on multimodal datasets to obtain scenario model tools (analogous to developing specialized weapons such as fighter jets, warships, and sniper rifles for different combat zones). Lightweight model tools refer to small, compressed, optimized, and automatically generated models that can run efficiently in resource-constrained environments. Without sacrificing accuracy, they significantly reduce the number of model parameters, computational cost (FLOPs), and storage size, thereby improving inference speed and reducing power consumption and latency. To ensure that scenario model tools possess low latency and efficient inference capabilities while meeting task accuracy requirements, all scenario model tools are trained and then centrally integrated to build a scenario model library (analogous to the basic weapon reserves in an arsenal). This solves the problems of large structure and high inference latency associated with traditional multimodal large models, providing adapted tools for industrial real-time scenarios. Based on the knowledge organization structure and scenario model library, the core parameters such as the input data format, type, and accuracy requirements of each scenario model tool are clearly defined, leading to the construction of a dedicated database. The core function of this database is to store various input data required for the operation of scenario model tools (such as real-time sensor data, historical process parameters, image and video data, analogous to the ammunition required for the use of stockpiled weapons), while simultaneously enabling efficient data retrieval, updating, and management. This ensures that scenario model tools can quickly acquire high-quality input data during operation, guaranteeing the stability and accuracy of model inference.Based on multimodal datasets, the basic multimodal large model is fine-tuned and distilled. Fine-tuning allows the multimodal large model to learn the professional knowledge and data characteristics of specific industrial scenarios, improving its understanding of industrial multimodal data. A multimodal large model (MLM) can be a large-scale neural network capable of simultaneously understanding, generating, and fusing multiple modalities such as text, images, audio, video, and sensor data. Distillation simplifies the multimodal large model structure, reducing its complexity while retaining core reasoning capabilities. The resulting backup inference engine is not directly used for regular real-time tasks but serves as a high-level support (analogous to a strategist), providing auxiliary reasoning when scenario model tools cannot handle complex chained reasoning or high-difficulty tasks, thus improving the overall system's ability to solve complex problems. For each industrial sub-scenario, textual modal data (such as process documents, operation manuals, fault reports, industry standards, etc.) is extracted from the multimodal dataset and used as training data to fine-tune a medium-sized language model. Small and medium-sized language models can be distinguished by the number of parameters. For example, a small language model has fewer than 1 byte (billion) of parameters, while a medium-sized language model has parameters in the range of [1 byte, 10 bytes]. Through fine-tuning, the medium-sized language model learns the specialized terminology, process specifications, and problem-solving experience of the corresponding scenario, forming a scenario-specific expert agent for each industrial sub-scenarios (analogous to training soldiers with specialized combat skills in each combat zone). This scenario-specific expert agent focuses on handling text-based tasks (such as problem consultation and rule interpretation) within the industrial sub-scenarios, enhancing the accurate output of professional knowledge and reducing the risk of model illusion.
[0014] Furthermore, textual modal data is extracted from the multimodal dataset, and a small language model is fine-tuned to obtain the routing central command (analogous to an officer responsible for overall coordination); its core functions include the following: 1. Receive user instructions (analogous to reconnaissance reports) and accurately identify key information such as task requirements and core parameters in the user instructions; 2. Based on the recognition results of user commands and combined with the characteristics of each industrial sub-scenarios, calculate and generate a set of confidence scores for the target industrial sub-scenarios (analogous assessment of intelligence credibility). 3. Based on the confidence set of the target industrial sub-scenarios, an adaptive routing decision algorithm (analogous to military strategy) is invoked to select the most suitable target model from the operational model library (a comprehensive resource pool integrating scenario model libraries, backup inference agents, and scenario expert agents). This drives the target model to process the task and generate the target answer. Furthermore, the adaptive routing decision algorithm can dynamically adjust the target model selection based on the scenario correlation strength (analogous to the strength of reinforcement support), ensuring the accuracy and efficiency of task processing and achieving the coordinated operation of the entire system.
[0015] This application's embodiments, through industrial scenario segmentation and targeted knowledge system construction, allow the model to overcome generalization limitations, achieve precise scenario adaptation, and meet the segmented needs of different industrial fields and the entire lifecycle, avoiding output results from deviating from industrial standards and improving practicality. Lightweight scenario model tools replace traditional multimodal large models, reducing inference latency and enabling rapid response to the needs of segmented industrial scenarios. The routing central command, combined with an adaptive routing decision algorithm, achieves rapid instruction recognition and accurate model selection, improving overall response efficiency and meeting the real-time requirements of industrial control loops. Scenario model tools, backup inference agents, and scenario expert intelligent agents provide multi-faceted collaborative support, ensuring efficient processing of conventional industrial segmented scenarios while addressing complex inference needs through backup inference agents, reducing the risk of illusion. Furthermore, the modular model construction method for industrial segmented scenarios allows for the integration into the existing system by only supplementing multimodal fusion data and training the corresponding model when adding new industrial segmented scenarios, improving the overall architecture's flexibility, scalability, and reusability.
[0016] In another embodiment of this disclosure, an adaptive routing decision algorithm is used to select target models from a combat-style model library based on the confidence set of the target industrial sub-scenarios, and then the target model is used to generate the target answer, including the following cases: Case 1: If the confidence level of the highest target industrial sub-scenario in the set of target industrial sub-scenario confidence levels is greater than or equal to the preset main scenario threshold, determine the first target scenario expert agent corresponding to the target industrial sub-scenario with the highest confidence level; use the first target scenario expert agent to analyze the user command, and determine the first target scenario model tool from the scenario model library based on the analysis results; generate the first target answer based on the first target scenario model tool and the input data of the corresponding first target scenario model tool in the database. Scenario 2: If the confidence level of the highest target industrial sub-scenario in the set of target industrial sub-scenario confidence levels is less than the preset main scenario threshold but greater than or equal to the preset minimum threshold, then based on the knowledge organization structure, determine the assistance scenario whose association strength with the target industrial sub-scenario corresponding to the highest target industrial sub-scenario confidence level is greater than or equal to the assistance threshold; based on the assistance scenario, determine the second target scenario expert agent corresponding to the assistance scenario; use the second target scenario expert agent to analyze the user instructions, and determine the second target scenario model tool from the scenario model library based on the analysis results; generate the second target answer based on the second target scenario model tool and the input data of the corresponding second target scenario model tool in the database. Case 3: If the confidence score of the highest target industrial sub-segment in the set of confidence scores is less than the preset minimum threshold, the backup inference engine is used to analyze the user's instructions and generate a third target answer based on the multimodal fusion data.
[0017] In this embodiment of the disclosure, based on the confidence set of the target industrial sub-scenarios, an adaptive routing decision algorithm is used to select suitable target models from the operational model library to generate the target answer. For example... Figure 2 As shown, after the routing central command receives the user instruction (S201), it performs industrial sub-scenario classification based on the user instruction (S202), obtaining the confidence set of all possible target industrial sub-scenarios (S203), expressed by the formula: ; in, Indicates user commands, This represents a preprocessing function that denoises and standardizes terminology for user commands. This indicates the central command structure for routing. This represents the set of confidence scores for a specific industrial scenario, where each individual in the set is... Each corresponds to a confidence level for all possible target industrial sub-scenarios. This represents the highest confidence level among the target industrial sub-scenarios in the confidence set. This indicates the target industrial sub-scenarios corresponding to the highest confidence level of the target industrial sub-scenarios.
[0018] The confidence score of the highest target industrial sub-scenario is compared with the preset main scenario threshold and the preset minimum threshold, and expressed as follows: ; in, This indicates the preset main scene threshold. This indicates the preset minimum threshold.
[0019] Determine whether the confidence level of the highest target industrial sub-scenario is greater than or equal to the preset main scenario threshold (S204). If yes, proceed to step S206. If no, determine whether the confidence level of the highest target industrial sub-scenario is less than the preset main scenario threshold and greater than or equal to the preset minimum threshold (S205). If yes, proceed to step S207. If no, proceed to step S208.
[0020] If the confidence level of the highest target industrial sub-scenario is greater than or equal to the preset main scenario threshold, the first target scenario expert agent corresponding to the target industrial sub-scenario with the highest confidence level is determined (S206), and the process proceeds to step S209. This first target scenario expert agent possesses specialized knowledge specific to the target industrial sub-scenario and can accurately interpret user commands. The first target scenario expert agent analyzes the user commands, and based on the analysis results, determines the first target scenario model tool from the scenario model library (S209). Based on the analysis results, the first target scenario expert agent retrieves the appropriate first target scenario model tool (such as a lightweight CNN model for welding defect detection) from the scenario model library. Based on the first target scenario model tool and the corresponding input data in the database, a first target answer is generated (S210). The input data required by the first target scenario model tool is extracted from the database, driving the model to run and generating the first target answer, ensuring the professionalism and real-time nature of the answer.
[0021] If the confidence level of the highest target industrial sub-scenario is less than the preset main scenario threshold but greater than or equal to the preset minimum threshold, the knowledge organization structure is queried to expand related scenarios and determine the assistance scenario (S207), proceeding to step S211. Based on the knowledge organization structure, assistance scenarios with a correlation strength greater than or equal to the assistance threshold corresponding to the highest target industrial sub-scenario confidence level are determined, expressed by the formula: ; in, This indicates the assistance scenario, and there can be multiple assistance scenarios. express With the The strength of association between sub-sectors of industrial applications This indicates the threshold for assistance. The knowledge organization structure includes the strength of correlation between different industrial sub-scenarios (analogous to reinforcements).
[0022] Based on the assistance scenario, determine the second target scenario expert agent corresponding to the assistance scenario (S211); use the second target scenario expert agent to analyze the user instructions, and determine the second target scenario model tool from the scenario model library based on the analysis results (S212); generate the second target answer based on the input data of the second target scenario model tool and the corresponding second target scenario model tool in the database (S213); for cases where there are multiple assistance scenarios, there can be multiple second target scenario expert agents, and thus multiple second target answers can be generated. Integrate the multiple second target answers to determine the final second target answer.
[0023] If the confidence level of the highest target industrial sub-scenario is less than a preset minimum threshold, a backup inference engine is used to analyze the user command (S208), proceeding to step S214. The routing central command determines that the current confidence level cannot match any industrial sub-scenario or related scenario, and directly calls the backup inference engine in the operational model library. Based on multimodal fusion data, a third target answer is generated (S214). The backup inference engine integrates multimodal fusion data from the industrial field, performs in-depth analysis and chain reasoning on the user command, and generates a fallback answer, i.e., the third target answer. This solves the problem of complex and fuzzy tasks that lightweight target scenario model tools cannot handle, ensuring the robustness of the entire system.
[0024] Based on confidence threshold processing logic, dedicated target model resources are matched to user commands with different levels of clarity, effectively solving the problem of low adaptability of general models and ensuring that the output results meet the professional needs of specific industrial scenarios. In high-confidence scenarios, lightweight scene expert agents and scene model tools are prioritized, avoiding the complex computational latency of large multimodal models and adapting to scenarios with high real-time requirements such as industrial production. In medium-confidence scenarios, existing resources from the assistance scenario are reused, eliminating the need to retrain the model and significantly shortening the response cycle. Only complex tasks in low-confidence scenarios are handled by backup inference agents. This maximizes resource utilization efficiency, improves inference reliability and robustness, and balances response efficiency with real-time performance.
[0025] In another embodiment of this disclosure, the following method is used: based on the confidence set of the target industrial sub-scenarios, an adaptive routing decision algorithm is employed to select target models from the operational model library, and the target models are used to generate target answers. The method further includes: If the generated target answer is either the first target answer or the second target answer, determine the confidence level of the first target answer or the second target answer, expressed by the formula: ; in, Indicates the confidence level of the target answer. This indicates the confidence level of the output from the routing central command. Indicates the first The confidence level of the output of an expert agent for a specific target scenario. Indicates the first The confidence level output by the target scene model tool. , , Indicates the confidence weight. This represents the total number of expert agents in the target scenario. This indicates the total number of target scene model tools; If the confidence level of the first target answer is less than a preset confidence threshold, a backup inference engine is used to analyze the user's instructions and, based on the input data of the corresponding first target scenario model tool in the database, a fourth target answer is generated; or, If the confidence level of the second target answer is less than the preset confidence threshold, the backup inference body is used to analyze the user's instructions and generate the fifth target answer based on the input data of the corresponding second target scenario model tool in the database.
[0026] In this embodiment, a quantitative evaluation of the generated first and second target answers is performed. First, the confidence level of the first or second target answer is calculated using a formula that integrates the confidence levels of the routing hub, expert agents, and scenario model tools across multiple dimensions. Then, based on a comparison of the calculated results with a preset confidence threshold, it is determined whether to activate a backup inference agent. For low-confidence target answers that do not reach the threshold, deep reasoning is performed using the original input data to generate the corresponding fourth or fifth target answer, ultimately ensuring the reliability and accuracy of the output answer. This reflects the reliability assessment of the routing central command's matching results between user commands and target industrial sub-scenarios. The routing central command performs semantic parsing of user commands and extracts semantic feature vectors. Feature vectors for the target scenario are extracted from feature vector libraries of various industrial sub-scenarios. By calculating the cosine similarity between the semantic feature vectors and the target scenario's feature vectors, the reliability of the matching results can be determined. . Reflecting the The accuracy of parsing user instructions and the reliability of requirement extraction by the target scenario expert agent are assessed. The target scenario expert agent decomposes user instructions, generates structured requirement description text, and calculates the degree of fit between the requirement description text and the knowledge organization structure, which is used as the confidence score output by the target scenario expert agent. Reflecting the The reliability of a target scene model tool in performing inference tasks based on input data can be assessed. The maximum predicted probability output by the target scene model tool (the maximum value from the probability distributions of each category output by the target scene model tool) can be extracted as the confidence level of the target scene model tool's output. This can be achieved through the formula: ; Determine the confidence level of the first or second target answer.
[0027] When the confidence level of the first or second target answer is less than a preset confidence threshold, a backup inference engine is used to analyze the user command. Unlike generating a third target answer based on multimodal fusion data, this method directly uses the input data from the corresponding target scenario model tool (first or second target scenario model) for deep analysis and inference to generate a fourth or fifth target answer. It directly reuses the input data from the original target scenario model tool in the database, eliminating the need for re-collection and processing of data. This ensures the optimization effect of the target answer while avoiding the efficiency loss caused by repeated input data processing. The backup inference engine is only activated for target answers with low confidence, avoiding the large model's regular consumption of computing power and balancing response efficiency with resource utilization costs. This ensures that the output target answer meets the stringent accuracy and safety requirements of industrial scenarios, reducing production accidents or quality losses caused by erroneous responses. It enhances the anti-interference capability and fault tolerance of the entire industrial large-scale model application system, ensuring stable operation in complex industrial environments.
[0028] In another embodiment of this disclosure, the method further includes: When the target answer is generated using the alternative reasoning body, the target answer generated using the alternative reasoning body will be integrated into the knowledge organization structure, and the alternative reasoning body will undergo RAG implicit evolution; the target answers generated using the alternative reasoning body include: the third target answer, the fourth target answer, and the fifth target answer; The target answer generated by the backup inference agent is used as the truth value, and the user instructions are used as training samples to train the scene model tool, the scene expert agent, and the routing central command agent.
[0029] In this embodiment, the target answer generated by the backup inference body is structurally decomposed to extract core knowledge elements. These elements may include industry terminology, process parameters, fault causes, solutions, and scenario association rules. Knowledge is then supplemented by integrating these elements into a knowledge organization structure. The extracted structured knowledge elements are integrated into the knowledge organization structure (such as a knowledge graph or knowledge base) to add new knowledge nodes or supplement attributes; for example, the solution attribute of a fault node is added to the knowledge base. The RAG implicit evolution of the backup inference body uses the updated knowledge organization structure as a retrieval library. When the backup inference body performs a reasoning task again, it retrieves this library in real time using RAG technology, integrating the retrieved accurate knowledge into the reasoning process. This improves reasoning ability (i.e., implicit evolution) without fine-tuning the model parameters of the backup inference body, ensuring that the output results are more aligned with industrial realities in subsequent complex task processing. The target answer generated by the backup inference body is used as the truth value, and the user instructions are used as training samples. Based on the task type of the scenario model tool, the training samples are categorized according to industrial sub-scenarios for supervised iterative training. For example, in a scenario model tool for weld defect detection, the defect image corresponding to the user command is used as input, and the defect type and level judgment results generated by the backup inference agent are used as the ground truth to optimize the feature extraction and classification accuracy of the scenario model tool. Through backpropagation, iterative training is carried out on the corresponding scenario expert agent and routing central command agent to improve the anti-interference capability and stable operation level of the entire intelligent decision-making system.
[0030] From a combat analogy perspective, this large-scale industrial model constitutes a highly efficient and collaborative digital army. Its combat advantages and collaborative logic are clearly discernible: the routing central command is like a "frontline officer," which, with its lightweight and agile core characteristics, can instantly analyze "reconnaissance reports" (user commands), accurately assess "intelligence credibility" (the set of confidence levels for the target industrial sub-scenarios), and formulate tactical decisions based on "military tactics" (adaptive routing decision algorithms). If the "battle situation is clear" (confidence levels are met), it directly commands "soldiers" (scenario expert agents) to rush to "sub-battlefields" (industrial sub-scenarios) such as water, land, and air to perform tasks; if "frontline forces are insufficient" (resources in a single scenario cannot meet the needs), it immediately calls for resources from related scenarios to coordinate operations based on "reinforcement support strength" (scenario correlation strength). Under the overall command of the "officer," the "soldiers" accurately retrieve the required "ammunition" (input data) from the "ammunition depot" (database) and select suitable "high-quality equipment" (such as corresponding scenario model tools like fighter jets, warships, and sniper rifles) from the "weapon arsenal" (model library) to efficiently complete professional tactical actions (task processing). The backup reasoning unit acts as a "wise strategist" in the rear. When the "battle situation is unclear, intelligence is questionable, or the front line is blocked" (the command scenario is ambiguous, the confidence level is low, or the initial plan fails), it relies on its profound multimodal knowledge base and refers to "historical combat review experience and combat maps" (knowledge organization structure) to conduct in-depth analysis and output decisive strategies (secondary optimization of the target answer). More importantly, this "digital army" has the ability to evolve after the battle. After completing each "combat mission" (processing user commands), it will simultaneously update the "combat map" (iterate the knowledge organization structure), replenish "ammunition reserves" (improve the database), upgrade "weapons and equipment" (optimize model tools), and refine "combat tactics" (iterate the adaptive routing decision algorithm), ultimately achieving continuous evolution and capability enhancement of the entire combat system.
[0031] In another embodiment of this disclosure, the multimodal dataset includes: a visual modality, a textual modality, and a temporal modality; wherein, the visual modality is data representing spatial information of industrial sub-scenarios; the textual modality is data representing semantic information of industrial sub-scenarios; and the temporal modality is data representing temporal information of industrial sub-scenarios. Based on the multimodal dataset, a knowledge organization structure is constructed for each industrial sub-scenarios, expressed by the formula: ; in, Represents the knowledge organization structure. and These represent indexes for specific industrial scenarios. Indicates the first Industrial-related sub-scenarios Indicates the first Industrial-related sub-scenarios Representing visual modality, Represents text modality, Represents timing modes, Indicates the first Industrial sub-scenarios and the first The strength of association between sub-sectors of industrial applications This indicates the total number of industrial sub-scenarios.
[0032] In this embodiment, a knowledge organization structure is constructed for each industrial sub-scenarios based on a multimodal dataset. The multimodal dataset includes visual modalities, textual modalities, and temporal modalities. For example, visual modalities include at least one of the following: visible light images, videos, infrared thermal imaging, X-ray images, and 3D point clouds. Textual modalities include at least one of the following: equipment manuals, standard operating procedures (SOPs), alarm logs, maintenance records, and expert knowledge. Temporal modalities include at least one of the following: numerical sensor data, acoustic waveforms, and vibration waveforms. This knowledge organization structure systematically integrates dispersed multimodal data with cross-scenario relationships, facilitating rapid knowledge retrieval and retrieval by subsequent models, reducing information redundancy during inference, and improving the efficiency of industrial scenario task processing.
[0033] In another embodiment of this disclosure, the first... Industrial sub-scenarios and the first The correlation strength of industrial-related sub-scenarios includes: Using text modalities from specific industrial scenarios as the corpus for specific industrial scenarios The corpus was segmented using a word segmentation tool to obtain a bag-of-words. ; Statistical analysis of each word segment in the bag of words In the corpus The frequency of a word is obtained by counting the number of times it appears in the word segmentation. Based on the word frequency of each segment, the relative frequency of each segment is determined, expressed by the formula: ; in, Indicates a participle, Indicates the first The relative frequency of word segmentation in industrial sub-scenarios Indicates the first Word frequency in industrial sub-segments Indicates the first Bag of words for specific industrial scenarios Indicates the first The sum of word frequencies of all words in the bag of words for industrial sub-segments.
[0034] Sort the relative frequencies of each word in descending order, for example, make satisfy: ; Indicates the length of the bag of words.
[0035] The initial number of high-frequency words is determined based on the relative frequencies after descending order and the preset cumulative frequency threshold, expressed by the following formula: ; in, Indicates the first The number of initial high-frequency words in industrial sub-scenarios. Indicates the search variable. Represents the sorting index variable. Indicates the sorting order is number 1. The relative frequency of word segmentation. Indicates the preset cumulative frequency threshold; The final set of high-frequency words is determined based on the initial number of high-frequency words and the preset upper limit threshold for high-frequency words. The formula is as follows: ; ; in, This indicates the final number of high-frequency words. This indicates the preset upper limit threshold for high-frequency word segmentation. Indicates the first The industrial sub-scenarios are ranked as follows: The participle of position; Indicates the first The final high-frequency word segmentation set for industrial sub-scenarios; Based on the preset stop word set, the stop weights of words in the final high-frequency word segmentation set are determined, expressed by the formula: ; in, This indicates the stop weight of the words in the final high-frequency word segmentation set. This represents a pre-defined set of stop words. Based on the inverse scene frequency of the words in the final high-frequency word segmentation set and the preset inverse document frequency threshold, the inverse scene weight of the words in the final high-frequency word segmentation set is determined, and the formula is expressed as follows: ; in, This represents the inverse scene weight of the word segmentation in the final high-frequency word segmentation set. This represents the activation function. This indicates the total number of industrial sub-scenarios. Indicates the existence of word segmentation A collection of industrial sub-scenarios Indicates the inverse document frequency threshold; The relative specificity score weight is determined based on the statistical frequency of word segmentation in the final high-frequency word segmentation set within the specific industrial scenarios. The formula is as follows: ; ; in, Indicates the relative specificity score weight. Indicates the participle in the first position The number of times data is collected for specific industrial scenarios. Indicates the participle in the first position The maximum number of statistical counts for industrial sub-scenarios Represents a positive smoothing factor. Indicates the penalty index for the versatility of different scenarios; The filtering weights are determined based on the stop weights of the words in the final high-frequency word segmentation set, the inverse scene weights of the words in the final high-frequency word segmentation set, and the relative specificity score weights. The formula is as follows: ; in, Indicates the first Filtering weights for specific industrial scenarios; The average vector of word segmentation context in the final high-frequency word segmentation set is determined using a context model, expressed by the formula: ; ; in, Indicates the first Average vector of word segmentation context in industrial sub-segments. Indicates including word segmentation A collection of text sequences Represents a text sequence. This represents the sentence vector output by the context model; Based on the relative specificity score weights, determine the first... Industrial sub-scenarios and the first The total weight of the industrial-related sub-scenarios is expressed by the formula: ; in, Indicates the first Industrial sub-scenarios and the first The total weight of industrial sub-scenarios Indicates the first The final high-frequency word segmentation set for industrial-related sub-scenarios. Indicates a participle, Indicates the first Filtering weights for specific industrial scenarios; According to the Industrial sub-scenarios and the first The total weight of the industrial sub-segmentation scenario and the average vector of the word segmentation context are used to determine the first... Industrial sub-scenarios and the first The correlation strength of industrial-related sub-scenarios is expressed by the formula: ; in, Indicates the first Average vector of word segmentation context in industrial sub-segments.
[0036] In another embodiment of this disclosure, the scene model tool includes at least one of the following: detection tools, prediction tools, evaluation tools, control tools, diagnostic tools, and generation tools; Among them, detection tools are used to identify anomalies in industrial scenarios and include at least one of the following: anomaly detection tools, defect segmentation tools, and safety and compliance detection tools. Predictive tools are used to predict trends in industrial scenarios and include at least one of the following: lifespan prediction tools, quality trend prediction tools, and demand prediction tools. Assessment tools, used to conduct risk assessments in industrial scenarios, include at least one of the following: health assessment tools, capacity assessment tools, and risk assessment tools; Control tools, used to implement decision-making in industrial scenarios, include at least one of the following: production planning and scheduling tools, process parameter optimization tools, and path planning tools; Diagnostic tools, used to trace the source of faults in industrial scenarios, include at least one of the following: fault location tools, root cause analysis tools, and expert-assisted decision-making tools; Generative tools are used to automatically generate content in industrial scenarios, including at least one of the following: report generation tools, data synthesis tools, and code generation tools.
[0037] In this embodiment, scene model tools are standardized and categorized according to their functional positioning, clarifying the core role and subclass scope of each type of tool. This allows for rapid matching of suitable scene model tools based on the task type corresponding to user commands, reducing the time cost of model selection and ensuring the real-time requirements of industrial scenarios. For each sub-scenarios in the industrial sector, lightweight model tools are trained using multimodal datasets to obtain scene model tools, and a scene model library is constructed based on these tools. . Tools for representing target scene models This indicates the input requirements for the target scene model tool. This indicates the output description of the target scene model tool. This describes the functionality of the target scene model tool. The target scene model tool can include traditional machine learning algorithms such as Support Vector Machines, Random Forests, Gradient Boosting Trees, K-Means, Hidden Markov Models, Gaussian Processes, Principal Component Analysis, and Association Rules; it can also include deep learning algorithms such as Convolutional Neural Networks, Recurrent Neural Networks, Transformers, Graph Neural Networks, Generative Adversarial Networks, Variational Autoencoders, and Diffusion Models.
[0038] In another embodiment of this disclosure, a backup inference body is obtained by fine-tuning the distillation of a large multimodal model based on a multimodal dataset, as expressed by the formula:
[0039] in, The visual modality data feature extraction model includes at least one of the following: convolutional neural network, visual transformer, and residual network. The text modality data feature extraction model includes at least one of the following: bidirectional transformer model, robustly optimized BERT pre-training method, and text transformer. The model for extracting features from temporal modal data includes at least one of the following: temporal convolutional network, long short-term memory network, or Transformer. This represents the extracted visual modality data features. This represents the extracted text modal data features. This represents the extracted temporal modality data features. This represents the aligned multimodal features. This represents the feature projection function of visual modality data, such as linear layers, 1×1 convolutional layers, multilayer perceptrons, and ViT mapping heads. This represents the feature projection function of text modality data, such as a linear layer or a BERT pooling layer. This represents the projection function of temporal modality data features, such as linear layers, LSTM fully connected layers, and Transformer fully connected layers. This indicates the fine-tuning method, such as full parameter fine-tuning, LoRA / QloRA low-rank adaptation, prompt fine-tuning, freeze fine-tuning, and instruction fine-tuning. This indicates the base of a multimodal large model. express The proportion, for example, 70%, This represents the fine-tuned multimodal large model. This refers to distillation methods, such as characteristic distillation, attention distillation, counter-distillation, self-distillation, and sparse distillation. This represents the multimodal large model after fine-tuning distillation. The parameter quantity representing the standby inference body. This indicates the preset threshold value for the first parameter.
[0040] Based on the text modalities in the multimodal dataset, a small language model is fine-tuned to obtain the routing central command structure, expressed by the formula: ; in, This represents the fine-tuned small-scale language model, i.e., the routing central command unit. These represent small language model bases, such as Qwen2-2.5B, Baichuan2-2.6B, GLM3-2.7B, MiniCPM-2.4B, Llama-2-2.7B, Gemma-2B, and Falcon-3B. This represents the parameter quantity of the routing central command body. This indicates the threshold value of the second parameter; For each specific industrial scenario, a medium-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a scenario expert agent for each scenario, expressed by the formula: ; in, Indicates the first Expert intelligent agents for specific industrial scenarios; Indicates the threshold of association strength. Indicates the first Industrial sub-scenarios and the first Features of text modal datasets for scenarios with association strength greater than or equal to an association strength threshold in industrial sub-scenarios. This represents the finely tuned medium-sized language model. This indicates a medium-sized language model base, such as Qwen2-7B, Baichuan2-7B, ChatGLM3-6B, GLM-4-9B, MiniCPM-7B, and Llama3-8B-Chinese. The number of parameters representing the scene expert agent. This represents the threshold value of the third parameter.
[0041] In another embodiment of this disclosure, a first target scene expert agent is used to analyze user commands, and a first target scene model tool is determined from a scene model library based on the analysis results, including: The first target scenario expert agent is used to analyze the user command and determine the first subtask corresponding to the user command; Based on the first subtask, the first candidate scene model tool is determined from the scene model library, expressed by the formula: ; in, This indicates the tool for representing the first candidate scene model. This indicates the matching rules between subtasks and scene model tools. Indicates the first subtask; Based on the functional description metadata of the first candidate scene model tool, the first target scene model tool is selected, as expressed by the formula: ; in, The tool for representing the first target scene model. This indicates semantic similarity calculation. Metadata describing the functionality of the first target scene model tool. Indicates the threshold for matching the functional description; A second-target scenario expert agent is used to analyze user commands, and based on the analysis results, a second-target scenario model tool is determined from the scenario model library, including: A second target scenario expert agent is used to analyze user commands and determine the second subtask corresponding to the user commands; Based on the second subtask, a second candidate scene model tool is determined from the scene model library, expressed by the formula: ; in, This indicates the tool for representing the second candidate scene model. This indicates the matching rules between subtasks and scene model tools. Indicates the second subtask; Based on the functional description metadata of the second candidate scene model tool, the second target scene model tool is selected, as expressed by the formula: ; in, Tools for representing second target scene models This indicates semantic similarity calculation. Metadata describing the functionality of the second target scene model tool. This indicates the threshold for matching the functional description.
[0042] In this embodiment of the disclosure, the target scene expert agent analyzes the user's instructions and decomposes them into sub-tasks. Candidate scene model tools are obtained through preset sub-task and scene model tool matching rules. For example, as shown in Table 2, Table 2 lists the sub-task and scene model tool matching rules.
[0043]
[0044] Table 2 Based on the functional description metadata of the candidate scenario model tools, the target scenario model tools are filtered. According to the instruction parsing and task decomposition of the scenario expert agent, it is determined whether the subtask is a composite task. If it is not a composite task, a single target scenario model tool is invoked; if it is a composite task, multiple target scenario model tools are invoked sequentially according to the planned task execution order.
[0045] like Figure 3 As shown, Figure 3 A flowchart of the tool for determining a target scene model provided in the embodiments of this disclosure includes: S301. The target scenario expert intelligent agent analyzes the user's instructions and determines the sub-tasks; S302. Based on the matching rules between subtasks and scene model tools, candidate scene model tools are obtained; S303, Metadata describing the functionality of the query candidate scenario model tool; S304. Determine whether the function description matches and identify the target scene model tool; if yes, proceed to step S305; otherwise, remove the candidate scene model tool and return to step S303. S305. Confirm whether the subtask is a composite task; if yes, proceed to step S306; if no, proceed to step S307. S306. Call multiple target scene model tools in sequence according to the execution order of the planned tasks; S307, Invoke a single target scene model tool; process ends.
[0046] By using subtask and scenario model tool matching rules, we ensure that the matched target scenario model tool is highly compatible with the task requirements, avoid processing errors caused by model mismatch, and adapt to the diverse and professional task requirements of industrial scenarios.
[0047] Based on the same disclosed concept, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the industrial large-scale model construction method as described in any of the above embodiments.
[0048] Through the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of this disclosure can be implemented in hardware or by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0049] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes in the drawings are not necessarily essential for implementing this disclosure.
[0050] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0051] The sequence numbers of the embodiments disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0052] Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
Claims
1. A method for constructing a large-scale industrial model, applied to the field of industrial knowledge question answering, characterized in that, include: Acquire multimodal fusion data in the industrial sector; The industrial sectors are classified into subdivided industrial scenarios. Based on the multimodal fusion data, determine the multimodal dataset for each industrial sub-scenarios; Based on the multimodal dataset, a knowledge organization structure is constructed for each industrial sub-scenarios; For each specific industrial scenario, a lightweight model tool is trained based on the multimodal dataset to obtain a scenario model tool; And construct a scene model library based on the scene model tool; Based on the knowledge organization structure and the scene model library, a database is constructed, which is used to store the input data of the scene model tools; Based on the multimodal dataset, fine-tune the distillation of the large multimodal model to obtain a backup inference body; For each industrial sub-scenarios, a medium-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a scenario expert agent for each industrial sub-scenarios. Based on the text modalities in the multimodal dataset, a small language model is fine-tuned to obtain a routing central command entity. The routing central command entity is used to identify received user instructions and determine the confidence set of the target industrial sub-scenarios. The confidence set of the target industrial sub-scenarios is used to select target models from the operational model library using an adaptive routing decision algorithm, and to generate target answers using the target models. The operational model library includes: a scenario model library, a backup inference entity, and a scenario expert intelligent agent.
2. The method as described in claim 1, characterized in that, Based on the confidence set of the target industrial sub-scenarios, an adaptive routing decision algorithm is used to select target models from the operational model library, and then the target model is used to generate the target answer: If the confidence level of the highest target industrial sub-scenario in the set of confidence levels of target industrial sub-scenarios is greater than or equal to the preset main scenario threshold, determine the first target scenario expert agent corresponding to the target industrial sub-scenario with the highest confidence level. The first target scene expert agent analyzes the user command, and determines the first target scene model tool from the scene model library based on the analysis results; the first target answer is generated based on the first target scene model tool and the input data of the corresponding first target scene model tool in the database. If the confidence level of the highest target industrial sub-scenario in the set of confidence levels of target industrial sub-scenarios is less than the preset main scenario threshold but greater than or equal to the preset minimum threshold, then based on the knowledge organization structure, an assistance scenario is determined that has a correlation strength with the target industrial sub-scenario corresponding to the highest confidence level of the target industrial sub-scenario that is greater than or equal to the assistance threshold. Based on the assistance scenario, determine the second target scenario expert agent corresponding to the assistance scenario; The user instructions are analyzed by the second target scene expert intelligent agent, and the second target scene model tool is determined from the scene model library based on the analysis results; The second target answer is generated based on the input data of the second target scene model tool and the corresponding second target scene model tool in the database; If the highest confidence level of the target industrial sub-scenarios in the set of confidence levels is less than a preset minimum threshold, a backup inference engine is used to analyze the user instruction and generate a third target answer based on the multimodal fusion data.
3. The method as described in claim 2, characterized in that, The method further includes: based on the confidence set of the target industrial sub-scenarios, using an adaptive routing decision algorithm to select target models from the operational model library, and using the target models to generate target answers; If the generated target answer is either the first target answer or the second target answer, determine the confidence level of the first target answer or the second target answer, expressed by the formula: ; in, Indicates the confidence level of the target answer. This indicates the confidence level of the output from the routing central command. Indicates the first The confidence level of the output of an expert agent for a specific target scenario. Indicates the first The confidence level output by the target scene model tool. , , Indicates the confidence weight. This represents the total number of expert agents in the target scenario. This indicates the total number of target scene model tools; If the confidence level of the first target answer is less than a preset confidence threshold, a backup inference engine is used to analyze the user instruction and, based on the input data of the corresponding first target scenario model tool in the database, a fourth target answer is generated; or, If the confidence level of the second target answer is less than the preset confidence threshold, the backup inference body is used to analyze the user instruction and generate the fifth target answer based on the input data of the corresponding second target scenario model tool in the database.
4. The method as described in claim 3, characterized in that, The method further includes: When the target answer is generated using the backup reasoning body, the target answer generated using the backup reasoning body is integrated into the knowledge organization structure, and the backup reasoning body undergoes RAG implicit evolution; the target answer generated using the backup reasoning body includes: the third target answer, the fourth target answer, and the fifth target answer; The target answer generated by the backup inference agent is used as the truth value, and the user instruction is used as the training sample to train the scene model tool, the scene expert agent, and the routing central command agent.
5. The method as described in claim 1, characterized in that, The multimodal dataset includes: visual modality, text modality, and temporal modality; wherein, the visual modality is data representing spatial information of industrial sub-scenarios; the text modality is data representing semantic information of industrial sub-scenarios; and the temporal modality is data representing temporal information of industrial sub-scenarios. The knowledge organization structure for each industrial sub-scenarios is constructed based on the multimodal dataset, and the formula is expressed as follows: ; in, Represents the knowledge organization structure. and These represent indexes for specific industrial scenarios. Indicates the first Industrial-related sub-scenarios Indicates the first Industrial-related sub-scenarios Representing visual modality, Represents text modality, Represents timing modes, Indicates the first Industrial sub-scenarios and the first The strength of association between sub-sectors of industrial applications This indicates the total number of industrial sub-scenarios.
6. The method as described in claim 5, characterized in that, The first is determined in the following manner Industrial sub-scenarios and the first The correlation strength of industrial-related sub-scenarios includes: Using text modalities of industrial sub-scenarios as the corpus of industrial sub-scenarios, the corpus is segmented using a word segmentation tool to obtain a bag-of-words; The word frequency of each word in the bag of words is obtained by counting the number of times it appears in the corpus. Based on the word frequency, the relative frequency of each word is determined, as expressed by the formula: ; in, Indicates a participle, Indicates the first The relative frequency of word segmentation in industrial sub-scenarios Indicates the first Word frequency in industrial sub-segments Indicates the first Bag of words for specific industrial scenarios Indicates the first The sum of word frequencies of all words in the bag-of-words for specific industrial scenarios; The relative frequencies of each word segment are sorted in descending order, and the initial number of high-frequency words is determined based on the descending relative frequencies and a preset cumulative frequency threshold. The formula is as follows: ; in, Indicates the first The number of initial high-frequency words in industrial sub-scenarios. Indicates the search variable. Represents the sorting index variable. Indicates the sorting order is number 1. The relative frequency of word segmentation. Indicates the preset cumulative frequency threshold; Based on the initial number of high-frequency words and the preset upper limit threshold for high-frequency words, the final set of high-frequency words is determined, expressed by the formula: ; ; in, This indicates the final number of high-frequency words. This indicates the preset upper limit threshold for high-frequency word segmentation. Indicates the first The industrial sub-scenarios are ranked as follows: The participle of position; Indicates the first The final high-frequency word segmentation set for industrial sub-scenarios; Based on the preset stop word set, the stop weights of words in the final high-frequency word segmentation set are determined, expressed by the formula: ; in, This indicates the stop weight of the words in the final high-frequency word segmentation set. This represents a pre-defined set of stop words. Based on the inverse scene frequency of the words in the final high-frequency word segmentation set and the preset inverse document frequency threshold, the inverse scene weight of the words in the final high-frequency word segmentation set is determined, and the formula is expressed as follows: ; in, This represents the inverse scene weight of the word segmentation in the final high-frequency word segmentation set. This represents the activation function. This indicates the total number of industrial sub-scenarios. Indicates the existence of word segmentation A collection of industrial sub-scenarios Indicates the inverse document frequency threshold; The relative specificity score weight is determined based on the statistical frequency of word segmentation in the final high-frequency word segmentation set within the specific industrial scenarios. The formula is as follows: ; ; in, Indicates the relative specificity score weight. Indicates the participle in the first position The number of times data is collected for specific industrial scenarios. Indicates the participle in the first position The maximum number of statistical counts for industrial sub-scenarios Represents a positive smoothing factor. Indicates the penalty index for the versatility of different scenarios; The filtering weights are determined based on the stop weights of the words in the final high-frequency word segmentation set, the inverse scene weights of the words in the final high-frequency word segmentation set, and the relative specificity score weights. The formula is as follows: ; in, Indicates the first Filtering weights for specific industrial scenarios; The average vector of word segmentation context in the final high-frequency word segmentation set is determined using a context model, expressed by the formula: ; ; in, Indicates the first Average vector of word segmentation context in industrial sub-segments. Indicates including word segmentation A collection of text sequences Represents a text sequence. This represents the sentence vector output by the context model; Based on the relative specificity score weights, determine the first... Industrial sub-scenarios and the first The total weight of the industrial-related sub-scenarios is expressed by the formula: ; in, Indicates the first Industrial sub-scenarios and the first The total weight of industrial sub-scenarios Indicates the first The final high-frequency word segmentation set for industrial-related sub-scenarios. Indicates a participle, Indicates the first Filtering weights for specific industrial scenarios; According to the Industrial sub-scenarios and the first The total weight of the industrial sub-segmentation scenario and the average vector of the word segmentation context are used to determine the first... Industrial sub-scenarios and the first The correlation strength of industrial-related sub-scenarios is expressed by the formula: ; in, Indicates the first Average vector of word segmentation context in industrial sub-segments.
7. The method as described in claim 5, characterized in that, The scenario model tools include at least one of the following: detection tools, prediction tools, evaluation tools, control tools, diagnostic tools, and generation tools; The aforementioned detection tools, used to identify anomalies in industrial scenarios, include at least one of the following: anomaly detection tools, defect segmentation tools, and safety compliance detection tools. The predictive tools are used to predict trends in industrial scenarios and include at least one of the following: lifespan prediction tools, quality trend prediction tools, and demand prediction tools. The assessment tools mentioned above are used to perform risk assessments in industrial scenarios and include at least one of the following: health assessment tools, capacity assessment tools, and risk assessment tools. The control tools are used to implement decision-making in industrial scenarios and include at least one of the following: production planning and scheduling tools, process parameter optimization tools, and path planning tools. The diagnostic tools are used to trace the source of faults in industrial scenarios and include at least one of the following: fault location tools, root cause analysis tools, and expert-assisted decision-making tools. The generation tools are used to automatically generate content in industrial scenarios, and include at least one of the following: report generation tools, data synthesis tools, and code generation tools.
8. The method as described in claim 5, characterized in that, The process involves fine-tuning and distilling the large multimodal model based on the multimodal dataset to obtain a backup inference body, expressed by the following formula: ; in, This represents a feature extraction model for visual modality data. This represents a text modal data feature extraction model. This represents a feature extraction model for time-series modal data. This represents the extracted visual modality data features. This represents the extracted text modal data features. This represents the extracted temporal modality data features. This represents the aligned multimodal features. Represents the projection function of visual modal data features. This represents the projection function of text modal data features. Represents the projection function of time-series modal data features. Indicates the fine-tuning method. This indicates the base of a multimodal large model. express proportion, This represents the fine-tuned multimodal large model. Indicates the distillation method. This represents the multimodal large model after fine-tuning distillation. The parameter quantity representing the standby inference body. This indicates the preset threshold value for the first parameter. The routing central command structure is obtained by fine-tuning a small language model based on the text modalities in the multimodal dataset, as expressed by the formula: ; in, This represents a small, fine-tuned language model. This represents a small language model base. This represents the parameter quantity of the routing central command body. This indicates the threshold value of the second parameter; For each specific industrial scenario, a medium-sized language model is fine-tuned based on the text modalities in the multimodal dataset to obtain a scenario expert agent for each scenario, expressed by the formula: ; in, Indicates the first Expert intelligent agents for specific industrial scenarios; Indicates the threshold of association strength. Indicates the first Industrial sub-scenarios and the first Features of text modal datasets for scenarios with association strength greater than or equal to an association strength threshold in industrial sub-scenarios. This represents the finely tuned medium-sized language model. This represents the base of a medium-sized language model. The number of parameters representing the scene expert agent. This represents the threshold value of the third parameter.
9. The method as described in claim 2, characterized in that, The step of analyzing the user instructions using the first target scene expert agent and determining the first target scene model tool from the scene model library based on the analysis results includes: The first target scene expert agent is used to analyze the user instruction and determine the first sub-task corresponding to the user instruction; Based on the first subtask, the first candidate scene model tool is determined from the scene model library, expressed by the formula: ; in, This indicates the tool for representing the first candidate scene model. This indicates the matching rules between subtasks and scene model tools. Indicates the first subtask; Based on the functional description metadata of the first candidate scene model tool, the first target scene model tool is selected, as expressed by the formula: ; in, The tool for representing the first target scene model. This indicates semantic similarity calculation. Metadata describing the functionality of the first target scene model tool. Indicates the threshold for matching the functional description; The step of using the second target scene expert agent to analyze the user instructions and determining the second target scene model tool from the scene model library based on the analysis results includes: The second target scenario expert agent is used to analyze the user instruction and determine the second sub-task corresponding to the user instruction; Based on the second subtask, a second candidate scene model tool is determined from the scene model library, expressed by the formula: ; in, This indicates the tool for representing the second candidate scene model. This indicates the matching rules between subtasks and scene model tools. Indicates the second subtask; Based on the functional description metadata of the second candidate scene model tool, the second target scene model tool is selected, as expressed by the formula: ; in, Tools for representing second target scene models This indicates semantic similarity calculation. Metadata describing the functionality of the second target scene model tool. This indicates the threshold for matching the functional description.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for constructing an industrial large model as described in any one of claims 1 to 9.