Method, apparatus, device, medium and product for processing large model
By constructing a three-layer model and a dynamic computing network, efficient collaboration and resource integration of large models are achieved, solving the problems of resource waste and low collaboration efficiency in the process of large model access and adaptation in existing technologies, and improving the utilization rate of computing resources and the accuracy of task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GRP FUJIAN CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot guarantee the requirements of high accuracy, low latency, and interpretability during the integration and adaptation of large models. Especially in scenarios involving the fusion of multiple knowledge domains, simple parameter tuning and data optimization cannot guarantee accurate reasoning and determination, resulting in wasted computing resources and low collaborative efficiency.
A three-layer model with interpretable logical nodes is constructed, which abstracts and decomposes the large model into professor, teacher and student layers. The computing power fragments are integrated through dynamic computing power network, a collaborative mechanism is established, and the model collaboration relationship and computing power integration effect are optimized.
It improves the utilization rate of computing resources, reduces the capacity of a single computing resource, and achieves efficient model collaboration and task execution, solving the problems of resource waste and low collaboration efficiency in traditional large model processing.
Smart Images

Figure CN122114162A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of large model technology, and in particular to a method, apparatus, device, medium and product for large model processing. Background Technology
[0002] Currently, with the implementation of AI+ technology, the execution logic of inference models has become a hot topic in the context of Deepseek's integration with various scenarios. Especially in the application scenarios of intelligent office and smart computing power, inference models not only need to meet the core requirements of high accuracy, low latency and interpretability, but also need to cope with challenges such as complex task collaboration, heterogeneous resource scheduling and dynamic scenario adaptation.
[0003] In current technology, the integration and adaptation of numerous models can only be achieved by using massive amounts of samples with both positive and negative samples to constrain and improve the model's capabilities. Simple parameter tuning cannot guarantee the inference and output results of large models, especially in scenarios involving the fusion of multiple knowledge domains, such as large models used to handle transmission capacity or computing power scheduling. Optimizing the model with massive amounts of data or simply adjusting the parameters cannot guarantee the accuracy of the final inference. At the same time, the training process will waste computing resources to a great extent. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, medium, and product for processing large models, enabling models in complex scenarios to possess better deep-level understanding capabilities.
[0005] A first aspect of this disclosure provides a method for processing large models, the method comprising: A three-layer model with interpretable logical nodes is constructed, and any target large model that can be accessed is abstracted and decomposed into computing power units that are adapted to the three-layer model. The three-layer model includes: a professor layer that realizes link architecture control and decision guidance, a teacher layer that realizes instruction conversion and transmission, and a student layer that realizes task response and implementation. Perform overall traffic graph generation and networking operations for the computing power network, and construct a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can adaptively adjust based on changes in computing power distribution. Based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer model, a suitable master model is selected, and a collaborative mechanism is established between the models at each layer, with the professor layer in charge, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer model and the dynamic computing network. Based on the aforementioned collaborative mechanism, iterative operations are performed during model training. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, the model collaboration relationship and the integration effect of large model computing power are optimized, thus completing the processing of the target large model.
[0006] A second aspect of this disclosure provides an apparatus for large model processing, characterized in that the apparatus comprises: The building module is configured to build a three-layer model of interpretable logical nodes, which abstracts and decomposes any target large model into computing power units that adapt to the three-layer model. The three-layer model includes: a professor layer that implements link architecture control and decision guidance, a teacher layer that implements instruction conversion and transmission, and a student layer that implements task response and implementation. The execution module is configured to perform the generation of the overall traffic graph of the computing power network and the networking operation, and to build a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can be adaptively adjusted based on changes in computing power distribution. The filtering module is configured to filter suitable master models based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer models, and establish a collaborative mechanism among the models at each layer, with the professor layer leading, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer models and the dynamic computing network. The optimization module is configured to perform iterative operations during model training based on the aforementioned collaborative mechanism. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, it optimizes the model collaboration relationship and the integration effect of large model computing power, thereby completing the processing of the target large model.
[0007] A third aspect of this disclosure provides an electronic device, including: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is used to execute the instructions to implement the above-described method.
[0008] A fourth aspect of this disclosure provides a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods described above.
[0009] A fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0010] The above-mentioned at least one technical solution adopted in the embodiments of this disclosure can achieve the following beneficial effects: the fragmented task integration technology based on large inference models, through real-time decision inference and dynamic model to build a dynamic large model computing power network, integrates the resources in the computing power network and manages them by the computing power network itself, and at the same time, each of the computing power networks automatically performs model training / inference, and obtains the final result through result integration, so that the large model can be spliced together from small models distributed in different computing power nodes, effectively improving the utilization rate of computing power resources and reducing the capacity of a single computing power resource. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a method for processing large models provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating yet another method for processing large models provided in this disclosure embodiment; Figure 3 This is a schematic diagram of the structure of a large model processing apparatus provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an exemplary computer system provided in an embodiment of this disclosure. Detailed Implementation
[0014] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0015] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0016] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0018] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0019] The following is combined Figures 1-5 The present disclosure describes the methods, apparatus, devices, media, and products for large model processing provided in the embodiments.
[0020] Figure 1 A flowchart illustrating a method for processing large models provided in this disclosure is shown below. Figure 1 As shown, a method for processing large models includes: S101. Construct a three-layer model of interpretable logical nodes, and abstract and decompose any target large model into computing power units that adapt to the three-layer model. The three-layer model includes: a professor layer that realizes link architecture control and decision guidance, a teacher layer that realizes instruction conversion and transmission, and a student layer that realizes task response and implementation. The large model is abstracted into three logical node models: the link architecture model (professor model), the instantaneous action execution model (teacher model), and the response landing student model (student model).
[0021] The professor layer includes two functionally differentiated professor models, each corresponding to a different execution node, used to collaboratively implement link architecture management, decision guidance, and suitability assessment functions.
[0022] The step of abstracting and decomposing any target large model into computing power units adapted to the three-layer model includes: Based on the computing power requirements, task types, and functions of each level of the three-layer model, the target large model is decomposed into multiple computing power units that can be allocated to each level of the model. The computing power units can be adapted to the task processing requirements of the corresponding level.
[0023] This three-tiered hierarchical architecture separates the responsibilities of control, transmission, and execution, avoiding the inefficiency and chaotic management caused by a single model level bearing all tasks, and improving the hierarchical and standardized processing of large models. The interpretable design makes the collaboration logic and task allocation of the three-layer model traceable, solving the problems of difficult fault location and directionless optimization caused by black-box collaboration in traditional large model processing. The abstract decomposition of the target large model and its adaptation to the three-layer model enables fine-grained splitting and allocation of large model computing power, avoiding waste of computing resources, while ensuring precise matching of large model tasks with the functions of each layer, improving task execution adaptability. The adaptability design for arbitrary access breaks the access restrictions of specific large models, improving the versatility of the method and adapting to target large models of different types and computing power requirements. The adaptation of computing power units and layers ensures that each layer model only processes the tasks it is adapted to, improving task response speed and processing accuracy.
[0024] S102. Execute the overall traffic graph generation and networking operation of the computing power network to construct a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can be adaptively adjusted based on changes in computing power distribution. This includes generating the overall traffic graph and networking operations for the computing power network, including: Based on a pre-defined undirected graph of the computing power network, execution nodes for two teaching models are selected. The undirected graph of the computing power network has computing power services as nodes, network bandwidth between nodes as side lengths, and an undirected cyclic structure. The execution nodes of the two teaching models run the corresponding teaching models respectively, generating two functionally differentiated flow vector graphs; Detect and handle node overlap issues between the two flow vector graphs, assign overlapping nodes to the flow vector graph with higher fit and update it; By merging the two updated flow vector graphs, the overall flow graph of the computing power network is obtained, thus completing the networking of the dynamic computing power network.
[0025] The method for handling the node overlap problem is as follows: quantify the degree of cooperative adaptation between each overlapping node and the two flow vector graphs. The degree of cooperative adaptation is determined based on the link association parameters and node feature differences between the overlapping node and the non-overlapping nodes in the corresponding flow vector graph. Each overlapping node is uniquely assigned to the flow vector graph with a higher degree of cooperative adaptation.
[0026] This step generates an overall traffic graph of the computing power network, enabling global visualization of computing power distribution and link status. This provides data support for subsequent computing power allocation and collaborative operation, avoiding blind allocation of computing power. The dynamic computing power network is specifically adapted to the needs of computing power fragment integration, solving the problem that traditional fixed computing power networks cannot adapt to the dynamic changes of computing power fragments in large models, and improving the flexibility of computing power fragment integration. It provides computing power support for the three-layer model, ensuring sufficient computing power when each level of model task is executed, avoiding task lag and iteration interruption due to insufficient computing power, and ensuring the continuity of large model processing. It adaptively adjusts based on changes in computing power distribution, realizing dynamic scheduling of computing power resources. When computing power is scarce or redundant in a certain area, it can automatically adapt and adjust, improving the utilization rate of computing power resources. The networking operation realizes the physical linkage between the computing power network and the three-layer model, laying the hardware foundation for subsequent collaboration between various levels and computing power transfer.
[0027] S103. Based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer model, select a suitable supervisor model, establish a collaborative mechanism among the models at each layer with the professor layer in charge, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer model and the dynamic computing network. The selected and adapted supervisor models include: Based on the topological characteristics of the dynamic computing network, the degree of fit between the two professor models and the dynamic computing network is calculated respectively. The professor model with a higher degree of fit is used as the computing fragmentation supervisor model, and the other professor model is used as the hardware fragmentation supervisor model.
[0028] This step combines the characteristics of the computing network topology with the three-layer model functionality to select the supervisor model, ensuring that the supervisor model's control capabilities and adaptability match the current computing network state. This avoids collaborative chaos and control failures caused by unreasonable supervisor model selection. It clarifies the collaborative logic of professor-led, teacher-transmitted, and student-executed collaboration, establishing a clear two-way collaborative link to solve the problems of irregular collaboration and inefficient instruction transmission at each level, thus improving hierarchical collaboration efficiency. It achieves collaborative linkage between the three-layer model and the dynamic computing network, deeply binding computing power scheduling with model collaboration. The dynamic adjustment of the computing network can synchronously adapt to model collaboration needs, and model collaboration needs can drive computing network optimization in reverse, forming a collaborative closed loop between computing power and models. Furthermore, the establishment of the collaborative mechanism allows each level of model and computing network to form an organic whole, avoiding resource consumption caused by independent operation and improving the overall synergy of large model processing. The design of the supervisor model selection provides a core decision-making body for subsequent suitability assessment and dynamic adjustment, ensuring that subsequent optimization operations have a clear guiding direction.
[0029] S104. Based on the aforementioned collaborative mechanism, perform iterative operations during model training. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, optimize the model collaboration relationship and the integration effect of large model computing power, and complete the processing of the target large model.
[0030] The inter-layer model collaboration mechanism includes: Clearly define the hierarchical responsibilities and powers of professors, teachers, and students, and establish a two-way collaborative link: professors issue control instructions to teachers, teachers transmit transformed instructions to students, students provide feedback on task execution status to teachers, and teachers report collaborative effects to professors.
[0031] The hierarchical sample processing and distribution includes: based on the functional differences of each level of the three-level model, standardized labeled samples are distributed to the corresponding level models according to preset rules; the professor level guides the sample processing rules; the teacher level transmits samples and processing instructions; and the student level executes sample processing and reasoning tasks.
[0032] The suitability assessment includes: after the full completion of a single round of sample training / inference tasks, the two teaching models respectively conduct a quantitative assessment of the suitability of the model collaboration effect and the computing power integration strategy, and count the suitability based on the assessment results for subsequent adjustment of the master model and optimization of collaboration parameters.
[0033] The dynamic adjustment includes: dynamically optimizing the number of student node grouping targets and the number of student layer models corresponding to each teacher layer model based on the dispersion of the suitability assessment results, while adjusting the allocation of supervisor roles for the two types of teaching models to achieve adaptive optimization of model collaboration relationships.
[0034] This step is based on a collaborative mechanism to conduct iterative operations, ensuring that the iteration process does not deviate from the hierarchical collaborative logic. This allows each iteration to specifically optimize the collaborative effect and improve the effectiveness of iterative optimization. Hierarchical sample processing and distribution are adapted to the functions of each level, avoiding low processing efficiency and sample waste caused by chaotic sample distribution. This ensures that sample processing is accurately matched with the tasks of each level, improving the relevance of model training. The suitability assessment enables quantitative judgment of the model collaborative effect and the computing power integration effect, solving the problem of optimization without a basis in traditional large model processing and providing a clear judgment standard for subsequent dynamic adjustments. The dynamic adjustment of model collaboration parameters can optimize the collaboration relationship based on the suitability assessment results, solving the problems of collaborative imbalance and poor computing power integration effect during iteration, and continuously improving the accuracy and efficiency of large model computing power integration. The entire iterative optimization process forms a closed loop of execution, evaluation, and adjustment, ensuring continuous improvement in the processing effect of large models, ultimately achieving efficient and accurate processing of the target large model, while reducing the cost of manual intervention.
[0035] In this embodiment of the disclosure, the fragmented task integration technology based on the large inference model constructs a dynamic large model computing power network through real-time decision inference and dynamic model construction. The resources in the integrated computing power network are managed by the computing power network itself. At the same time, each model automatically performs model training / inference in the computing power network, and the final result is obtained through result integration. This allows the large model to be spliced together from small models distributed on different computing power nodes, effectively improving the utilization rate of computing power resources and reducing the capacity of a single computing power resource.
[0036] Figure 2 A flowchart illustrating another method for processing large models provided in this disclosure is shown below. Figure 2 As shown, the method includes: S200. Construct a three-layer model with interpretable logical nodes, and abstract and decompose any target large model into computing power units that adapt to the three-layer model. The three-layer model includes: a professor layer that implements link architecture control and decision guidance, a teacher layer that implements instruction conversion and transmission, and a student layer that implements task response and implementation.
[0037] The professor layer includes two functionally differentiated professor models: a computing power professor model and a computing power professor model. Coordinating teaching model with computing power .
[0038] S201, Generate a traffic flow diagram to guide the overall development of the computing power network; In some embodiments, the computing power teaching model Coordinating teaching models with computing power The nodes and links in the full set of computing power network models are analyzed, and new path flow vectors for the transmission of computing power network traffic are generated for each. Then, the computing power teaching model is applied. Coordinating teaching models with computing power The path flow vectors output by these two professor models are heterogeneously concatenated to obtain the overall guidance and propulsion flow graph of the computing power network. .
[0039] S2011, Preliminary selection of execution nodes for the dual-professor model; This step specifically involves: selecting the computing power professor model. execution node And the computing power coordination professor model execution node .
[0040] Computing power network management node First, select nodes with high idle time in the current computing network as management nodes for the teaching model. The specific selection method is as follows: Computing power network management node Run the professor's model inference node selection algorithm, which uses an undirected graph computing network. As input, the node capability management interactive computing power network supports the strength reduction method, and two nodes are obtained from the undirected graph nodes of the network, which serve as the execution nodes of the computing power teaching model. Execution nodes of the teaching model coordinated with computing power .
[0041] Among them, the computing power network undirected graph It is an undirected cyclic graph constructed with computing services in the computing power network as nodes and the network bandwidth between nodes as the side length.
[0042] This step filters candidate nodes with high idle time and suitability for the dual-professor model from the computing power network, laying the foundation for subsequent precise selection of execution nodes for the computing power-professor model. Execution nodes of the teaching model coordinated with computing power The scope was defined, prioritizing nodes with high idle time in the computing network to ensure sufficient resources for the dual-professor model to run. An undirected graph of the computing network was also defined. The construction rules.
[0043] S2012, Professor's algorithm for initial selection of inference nodes in the model; First, load the undirected graph of the computing power network. Traversing the undirected graph of the computing power network All vertices, for any vertex Get all its edges For any edge When the side length value If the edge length is not updated using the algorithm for filtering inference nodes in the teaching model, then the update should be performed as follows:
[0044] in, , , All are uniform adjustment coefficients. Represents a node Currently available computing power Represents a node Currently available networks Represents a node Currently available storage.
[0045] The subscript 'c' has no actual meaning; it is mainly used to distinguish between the representation of the computing power teaching model and the computing power coordination teaching model, one being 'c' and the other 'ca'.
[0046] The .cn symbol is an identifier and has no actual meaning.
[0047] The subscript i represents the i-th node and the j-th node in the computing power network. Since there are edges between nodes, i and j are used to distinguish the two nodes.
[0048] If the edge length has already been updated by the edge length filtering algorithm of the teaching model inference node, then skip this update.
[0049] This step involves the undirected graph of the computing power network. The edge length value is updated in a personalized manner, incorporating core available resource indicators such as computing power, network, and storage into the edge length calculation. This upgrades the edge length from a simple network bandwidth indicator to a quantitative indicator that comprehensively reflects the node's resource capabilities and link bandwidth. This provides a basis for subsequent calculations of edge lengths and precise selection of execution nodes for the computing power teaching model. Execution nodes of the teaching model that coordinate computing power Provide quantitative evidence (after the edge length is updated, the total edge length of the node can directly reflect its adaptability as an execution node in the dual-professor model).
[0050] S2013, precise selection of execution nodes for the dual-professor model; Specifically: computing power network management node Calculate each node Summary of side lengths Select the node with the largest summation edge length as the computing power teaching model. execution node The node with the second longest summative edge length after the node with the longest summative edge length is selected as the computing power coordination teaching model. execution node .
[0051] This step calculates the total edge length of each node based on the updated edge length values in S2012, and sorts them from largest to smallest to accurately select the execution node for the computing power teaching model. Execution nodes of the teaching model that coordinate computing power .
[0052] The final determination of the execution node for the dual-professor model is achieved through quantization and sorting, ensuring the selection of the execution node for the chosen computing power professor model. Execution nodes of the teaching model that coordinate computing power These are the two nodes with the best overall resource capabilities in the computing power network, providing the optimal hardware platform for the subsequent stable operation of the dual-professor model and for streamlining the nodes / links of the computing power network.
[0053] S2014, the dual-professor model generates its own flow vector graph; Specifically: Execution Node In the middle, using computing power network undirected graph As input, run the computing power teaching model The computing power network flow vector graph is obtained. At the execution node In the same way, undirected graphs in computing power networks are used. As input, execute the computational power coordination teaching model. The computational power coordination network flow vector graph is obtained. .
[0054] Computational power professor model Based on the model's inherent organizational capabilities, this model can subsequently utilize any existing large-scale model capable of performing network structure organization. Furthermore, leveraging the capabilities of these existing large-scale models, it will organize the vertices, edges, and structure of the undirected graph in the computing power network to ensure a better match between the nodes and edges in the computing power network and the taught model. This step is a black-box process.
[0055] In this step, the selected dual-professor model execution node is enabled. , Run the corresponding professor model respectively, using the computing power network undirected graph. Each input generates its own flow vector graph. , These two flow vector diagrams serve as the overall guiding flow diagram for subsequent heterogeneous splicing. The core foundational data, while also allowing the dual-professor model to complete the initial sorting of the computing power network, ensuring that the network nodes / links match the model's operational requirements.
[0056] S2015, Dual-model flow vector graph node set overlap processing; Specifically, it refers to: the computing power network flow vector graph. Node set Coordinating network flow vector graph with computing power Node set Overlap (i.e.) When performing the overlapping load balancing process, the computing power network flow vector graph is processed. Coordinating network flow vector graph with computing power Update the computing power network flow vector graph. Node set Coordinating network flow vector graph with computing power Node set Non-overlapping (i.e.) If the condition is met, proceed directly to the next step.
[0057] The specific execution process of the management overlap balancing process is as follows: Traverse overlapping nodes Calculate overlapping nodes Computational power network flow vector graph Synergistic integration affinity :
[0058] in, Represents a node With nodes The side length is 0 when it is unreachable. Represents a node With nodes The characteristic differences are calculated using the following logic:
[0059] Recalculate overlapping nodes Coordinating network flow vector graph with computing power Synergistic integration affinity :
[0060] when When overlapping nodes Distribution to computing power network flow vector graph In the middle, otherwise overlapping nodes will be used. Distribution to the computing power coordination network flow vector graph middle.
[0061] In this step, the computing power network flow vector graph is detected. Node set Coordinating network flow vector graph with computing power Node set If overlapping nodes exist, the overlapping balance management process is executed to update the two vector graphs and resolve the node resource conflict problem (if the same node is occupied by two models at the same time, it will lead to insufficient resources and model lag). If no overlapping nodes exist, the process proceeds directly to the next step to ensure the independence of the node sets of the two vector graphs, thereby providing conflict-free basic data for subsequent heterogeneous splicing.
[0062] S2016, Generate the final overall guidance flow diagram for the computing power network. .
[0063] Final computing power network management node Directing computing power network flow to vector graph Coordinating network flow vector graph with computing power Merging the data yields an overall guidance for the advancement of the computing power network's traffic graph. And distribute it to the computing power teaching model. execution node Coordinating teaching models with computing power execution node .
[0064] S2011-S2015 ultimately output a conflict-free, two-dimensional complementary computing power network flow vector graph. (Computing power scheduling dimension) and computing power coordination network flow vector graph (Computing power coordination dimension) The two vector graphs can only represent the single-dimensional features of the computing power network and cannot support the subsequent integration of computing power fragments throughout the entire process.
[0065] This step merges the two-dimensional vector graphs into a unified guidance flow graph for the computing power network through a merging operation. It achieves unified integration of topological features in two dimensions: computing power scheduling and computing power coordination, generating a standardized core topology benchmark for the entire computing power network. This provides a unified core data basis for all subsequent networking, integration, and iteration operations, avoiding operational deviations caused by multi-dimensional data chaos.
[0066] The merged overall flow chart Distributed to dedicated execution nodes of the dual-professor model , Let computing power teach the model Computational power coordination teaching model By sharing the same set of computing power network topology benchmarks, data alignment and cognitive unification between the two models can be achieved.
[0067] S202, quantitative screening of the computing power fragmentation / hardware fragmentation supervisor's model; Each professor advances the flow chart according to the overall guidance of the computing power network, following the execution nodes. Flow vector diagram in and Calculate the deviation from its initial flow direction vector. The model with smaller deviations was used as the supervising professor model for computing power fragmentation, and the other was used as the supervising professor model for hardware fragmentation.
[0068] Computational power professor model execution node Calculate the overall network traffic graph Computational power network flow vector graph Deviation of the initial flow direction vector :
[0069] in, Represents the flow vector graph of computing power network The total edge length of the out-degree of the i-th node. Represents the flow vector graph of computing power network The total edge length of the in-degree of the i-th node. Represents the network flow vector graph Total number of nodes Represents the overall network traffic graph The total edge length of the out-degree of the i-th node. Represents the overall network traffic graph The total edge length of the in-degree of the i-th node. Represents the overall network traffic graph The total number of nodes.
[0070] Similarly, the computing power coordination teaching model execution node Calculate the overall network traffic graph using the following method. Coordinating network flow vector graph with computing power deviation :
[0071] in, Represents the flow vector graph of the computing power coordination network. The total edge length of the out-degree of the i-th node. Represents the flow vector graph of the computing power coordination network. The total edge length of the in-degree of the i-th node. Represents the flow vector graph of the computing power coordination network. The total number of nodes.
[0072] when At that time, the computing power teaching model will be used. As the supervising professor model for computing power fragmentation nodes, the computing power coordination professor model The supervising professor model as a hard-fragmented node. When At that time, the computing power will be coordinated to teach the model. As the supervising professor model for computing power fragmentation nodes, the computing power professor model The supervising professor model as a hard-fragmented node.
[0073] This step uses a deviation quantification formula based on the overall flow chart. The topological characteristics define the dedicated supervisors for computing power fragments / hardware fragments in the dual-professor model, enabling quantitative and precise decision-making on model division of labor. At the same time, it clarifies the core decision-making body for the three-layer model collaboration, laying the top-level division of labor rules and decision-making foundation for the specific implementation of subsequent computing power fragment integration.
[0074] S203, the supervising professor model of computing power fragment nodes, confirm the gradient order of the teacher model, and confirm the teacher rank group belonging to different professor models; This step is led by the supervising professor model of the computing power fragment nodes. It confirms the gradient order of the teacher model and divides the teacher ranking groups under different professor models. This establishes a hierarchical management relationship from the professor level to the teacher level in the three-layer model, clarifying the grouping, gradient, and affiliated management body of teacher nodes. This lays the foundation for the management and collaboration of the teacher level in subsequent instruction transmission from the teacher level to the student level and the hierarchical integration of computing power fragments / hardware fragments. Specifically, it includes the following steps: S2031, Node Classification and Flow Graph Segmentation; The supervisory professor model for fragmented computing nodes executes the node classification process and guides the overall network towards the traffic graph. Delete execution node With execution node Post-professorless network overall traffic graph By using a node segmentation method, the overall traffic graph of the unprofessional network is obtained. The overall traffic graph of the teacher node network is divided into segments. Overall traffic graph of student node network .
[0075] The specific execution flow of the node segmentation method is as follows: First, analyze the overall network traffic diagram without a professor. Traverse all nodes in the array and calculate the result when traversing to any node. The joint degree of nodes at time :
[0076] in, Represents a node To the node The side length, Represents a node The sum of the lengths of all out-degree sides, Represents a node The sum of the lengths of all in-degree sides, Represents a node The sum of the lengths of all sides with in and out degrees.
[0077] Take the joint global degree of the nodes The node at its maximum value is used as the teacher node. And without teaching the overall network traffic graph Extract all teacher nodes The edges between them form the overall flow graph of the teacher node network. The remaining nodes serve as student nodes. Similarly, there was no professor teaching the overall network traffic graph. Extract all student nodes The edges between them form the overall traffic graph of the student node network. .
[0078] Overall connectivity refers to the connection relationship between a node and all other nodes. The node with the highest overall connectivity, along with all nodes related to it, forms a teacher node. In other words, the node with the highest overall connectivity and all nodes directly connected to it are all considered teacher nodes.
[0079] S2032, Teacher node gradient grouping; The supervising professor model for computing power fragmentation nodes executes a gradient pairing process for teachers, based on the overall network traffic graph of the teacher nodes. All teacher nodes Grouping, resulting in teacher node groups .
[0080] Traverse teacher nodes For any ungrouped i-th teacher node Iterate through and calculate its relationship with other ungrouped teacher nodes. The combined profit and loss of the same group :
[0081] in, The calculation method for computing power balance is as follows:
[0082] in, For nodes The computational capability vector, For nodes The computational capability vector. This indicates taking the maximum value between x and y.
[0083] This refers to bandwidth utility. It is derived from the amount of data that can be transmitted within the statistical delay reception range.
[0084] The calculation method for the delay penalty is as follows:
[0085] in, For nodes With nodes Average network bandwidth latency between This represents the maximum acceptable delay.
[0086] When the joint profit and loss of the same group Greater than the lower limit of joint income (At the preset value) ungrouped teacher nodes Included in teacher nodes In the temporary group, traverse its nodes and other ungrouped teacher nodes. After completion, the teacher node will be... Integrating into teacher nodes Ungrouped teacher nodes in the temporary group are grouped into the same group.
[0087] It is understandable that teachers are grouped at different levels. It can adapt to different types and difficulties of computing power fragment integration tasks, providing a hierarchical execution entity for professors to issue differentiated integration instructions to teachers, thereby improving the efficiency and accuracy of fragment integration.
[0088] S2033, Teacher Grouping and Assignment; The supervising professor model for computing power fragmentation nodes groups the teacher nodes. The management is divided into groups, with teachers assigned to the corresponding computing power fragment nodes within the subjective teaching model's management scope. The specific division method is as follows: Traverse teacher node groups For any i-th teacher node group Calculate the computing performance bias of this group. :
[0089] in, Indicates teacher node grouping Number of nodes This indicates the number of teacher nodes.
[0090] When performance bias Greater than the average node performance If the group is in a certain condition, it will be assigned to the subjective teaching model management scope of the computing power fragment node; otherwise, it will be assigned to the subjective teaching model management scope of the hard fragment node.
[0091] This step establishes a dedicated collaborative link between the dual-professor model and the teacher layer. Specifically, it establishes dedicated instruction transmission and feedback links between the computing power fragment / hardware fragmentation supervisor professor model and the corresponding teacher group. Subsequently, the dual-professor model can directly issue integration instructions to the teacher groups within its own management scope, and the teacher groups can also directly report the execution status to their respective professor models, thereby improving collaborative efficiency.
[0092] S2034. Teacher group data is synchronized to the hard-segmented supervisor professor model; Teacher node grouping After the division is completed, the divided teacher nodes will be grouped. With student nodes Overall traffic graph of student node network The supervising professor model is sent to the hard fragment node to achieve full synchronization of the dual-professor model in terms of data at the teacher and student levels.
[0093] S204, The supervising professor model of the hard-broken node confirms the student nodes that the student model depends on during its operation. Grouping And the number of teachers belonging to different teacher models; This step includes: S2041, the supervising professor model of hard-fragmented nodes, calculates student nodes. Grouping Target quantity :
[0094] in, Indicates the number of teacher nodes. This indicates the number of student nodes. This means that when the value of y is less than x, we take x; when the value of y is greater than z, we take z; otherwise, we take y.
[0095] The principal professor model of hard-fragmented nodes starts from student nodes. Random selection Given a set of student nodes as the initial nodes, the k-nearest neighbor algorithm is executed. During the execution of the k-nearest neighbor algorithm, the student nodes... With student nodes The distance calculation method is as follows:
[0096] The subscript i represents the i-th node and the j-th node in the computing power network. Since there are edge relationships between nodes, i and j are used to distinguish the differences between two nodes.
[0097] in, Represents student nodes The model's computational capability vector. Represents student nodes With student nodes Network bandwidth between.
[0098] Finally, the professor model that yields hard-fragmented nodes first calculates the student nodes. Grouping .
[0099] This step provides a quantitative basis for grouping student nodes, based on the target number. Clearly define the total number of student groups needed to avoid imbalances in teacher-student collaboration caused by too many or too few student groups, ensuring that the number of student groups and teacher groups are appropriately matched; and match the actual needs of fragment integration with the target number. The computation combines the capabilities of the teaching staff with the need for fragment integration, matching the size of student groups with the actual task of integrating fragmented computing power / hardware fragments, thus avoiding resource waste or insufficient capabilities.
[0100] S2042, Quantify the progressive efficiency of the calculation model, complete the group association of teachers and students, and determine the number of affiliations; Traverse student nodes Grouping For the i-th student node group Traverse teacher node groups For the j-th teacher node group Student node grouping Associated with teacher node group Model progression efficiency The calculation method is as follows:
[0101] in, Indicates teacher grouping The number of associated student groups Indicates teacher grouping computing power Indicates teacher grouping Non-computing hardware capabilities, This represents the average computing power of student nodes. This represents the average non-computing hardware capability of student nodes. Student node grouping Quantity, This represents the variance of x and y.
[0102] After traversing all teacher node groups, select the teacher node group with the highest progressive efficiency between the student group node and the teacher group model as its associated teacher group, and increment the number of associated student groups of this teacher group by 1.
[0103] S205. Based on the given labeled samples, perform inference iterations to obtain inference fragments, and generate an integration strategy for the inference fragments according to their inference progress. During each round of training / derivation, given The labeled samples are used for one round of inference iteration to obtain inference fragments. Based on the inference progress of the inference fragments, an integration strategy for the inference fragments is generated. During each round of training / derivation, obtain from the sample dataset For each unlabeled sample, the teacher model sends training / inference samples in batches through the professor model. After receiving the batched training / inference samples, the teacher model groups one unlabeled sample for each student node and distributes them for training / inference one by one. Each student node has its corresponding training / inference. After the student node completes the training / inference, the result is returned to the corresponding teacher model. The teacher model then reissues an unexecuted sample from the unexecuted sample list until the teacher model determines that all sample training / inference has been completed. Finally, the results are handed over to the professor model to summarize the final results.
[0104] The professor model negotiates and assigns samples based on the type of the supervising professor model and the combined capabilities of the predicted samples to be labeled and the managed models. During the negotiation process, the samples to be labeled can be automatically assigned to a management node with available resources if there are available resources, and otherwise assigned to a professor model with the corresponding bias if there are no available resources.
[0105] After all labeled sample inferences are completed, the computational power teaching model... Coordinating teaching models with computing power Based on the student model, associated teacher model, and professor model that perform inference based on labeled samples, the integration strategy for each labeled sample is calculated after each labeled sample training / inference progress report. :
[0106] in, This indicates the model size for inference with labeled samples. This represents the computational power of the student model performing inference on labeled samples. To achieve the required proportion, The execution time is [time already elapsed]. .
[0107] When integration strategy At that time, the proportions are combined, that is, the final result. The update method is as follows:
[0108] When integration strategy At that time, the complete merge, that is, the final result. The update method is as follows:
[0109] in, This represents the intermediate results when reporting the training / inference progress for each labeled sample. This indicates the percentage of completion when reporting the training / inference progress for each labeled sample.
[0110] S206, Computing Power Professor Model Coordinating teaching models with computing power The suitability of the strategy is evaluated, and the one with high suitability is the supervisory professor model for the next round of computing power fragmentation; Computational power professor model Coordinating teaching models with computing power The suitability of each calculation strategy, and the percentage of completion when progress feedback is received. At that time, determine the appropriate strategy The type of integration strategy At that time, the computing power teaching model The degree of fit increments by 1; otherwise, the computational power coordination professor model... The degree of suitability increases by 1.
[0111] Once all labeled samples have been trained / inferred, compare the computational power of the teaching model. Coordinating teaching models with computing power The degree of suitability is determined by taking the model with high suitability as the master professor model for the next round of computing power fragmentation, and the model with low suitability as the master professor model for the next round of non-computing power fragmentation.
[0112] S207. Repeat S203-S205, and simultaneously in S204 and S205, increase the number of student models belonging to different teacher models in S204 by adjusting the appropriateness variance. Repeat steps S203-S205, and simultaneously add adjustments to student nodes using appropriateness variance in steps S204 and S205. Grouping Target quantity :
[0113] in, Representing the computing power professor model Coordinating teaching models with computing power The variance of the degree of fit.
[0114] Then, based on the remaining steps in S204, the number of student models belonging to different teacher models is automatically adjusted.
[0115] S208. Re-evaluate the suitability of the model logic network. If the integrated fragment volume is less than the compressibility volume of the inference volume, the integrated inference is completed; otherwise, execute S206 and S207.
[0116] After iteration, the degree of the computational model logic network is determined. :
[0117] in, This represents the total number of progress reports for labeled samples in the current iteration. This is the total running time of the logic network in this iteration.
[0118] At this point, the total volume of the integrated fragments is...
[0119]
[0120] in, This represents the volume of the lower j labeled samples.
[0121] Target reasoning volume compression acceptance volume :
[0122] in, This is for the acceptance of compression of inference volume.
[0123] when If the integration reasoning is completed, the process ends; otherwise, execute S206 and S207 for another iteration.
[0124] Furthermore, an early stopping mechanism is employed during model training to prevent overfitting. This means that if the validation set loss does not show significant improvement within several iterations while the training set loss continues to decrease, the training process is terminated early. This ensures that the model maintains its ability to fit the training data while possessing good generalization performance. Through these systematic model training and optimization strategies, the constructed model can accurately classify newly generated logs. It can not only accurately identify logs of different levels such as normal, warning, and error, but also further subdivide and classify transaction logs, error logs, and slow query logs, providing a solid and reliable basis for subsequent log analysis, fault diagnosis, and database performance optimization.
[0125] In this embodiment, a large inference model is provided for integrating storage and computing fragment tasks in a computing power network server in a smart office scenario. This model uses real-time decision-making inference and dynamic model building to construct a dynamic large-scale computing power network. Resources within the integrated computing power network are self-managed by the network itself. Simultaneously, model training / inference is automatically performed within the network, and the final result is obtained through result integration. This allows the large model to be constructed by splicing together smaller models distributed across different computing power nodes, effectively improving computing power resource utilization and reducing the capacity of a single computing power resource. The provided method for generating an overall computing power network traffic graph and forming a computing network uses a multi-layered approach. The dynamic networking method divides the network into three layers: professor, teacher, and student models. Based on the network conditions, it completes self-organizing networking, dynamically adjusting itself during the process. This avoids the limitations of fixed-rule networking methods, and its layered networking mode facilitates unified management. The iterative method provided for large model training / derivation employs a strategy of dynamically adjusting the computing network during iteration, ensuring good scalability of the entire computing network during task execution. Furthermore, based on the scale and suitability of the results, it dynamically adjusts the roles and tasks of each participating node, further improving the accuracy of allocation strategies during task execution.
[0126] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements, optimizations and modifications can be made without departing from the principle of the present invention, and these should also be considered within the scope of protection of the present invention.
[0127] Figure 3 This is a schematic diagram of the structure of a large model processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, the device 300 includes: The construction module 301 is configured to construct a three-layer model of interpretable logical nodes, which abstracts and decomposes any target large model into computing power units that adapt to the three-layer model. The three-layer model includes: a professor layer that implements link architecture control and decision guidance, a teacher layer that implements instruction conversion and transmission, and a student layer that implements task response and implementation. The execution module 302 is configured to perform the generation of the overall traffic graph of the computing power network and the networking operation, and to build a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can be adaptively adjusted based on changes in computing power distribution. The filtering module 303 is configured to filter suitable master models based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer models, and establish a collaborative mechanism among the models at each layer, with the professor layer in charge, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer models and the dynamic computing network. The optimization module 304 is configured to perform iterative operations during model training based on the aforementioned collaborative mechanism. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, it optimizes the model collaboration relationship and the integration effect of large model computing power, thereby completing the processing of the target large model.
[0128] In some embodiments, the professor layer includes two functionally differentiated professor models, each corresponding to a different execution node, for collaboratively implementing link architecture management, decision guidance, and suitability assessment functions.
[0129] In some embodiments, the step of abstracting and decomposing any access-based target large model into computing power units adapted to the three-layer model includes: Based on the computing power requirements, task types, and functions of each level of the three-layer model, the target large model is decomposed into multiple computing power units that can be allocated to each level of the model. The computing power units can be adapted to the task processing requirements of the corresponding level.
[0130] In some embodiments, performing overall traffic graph generation and networking operations for the computing power network includes: Based on a pre-defined undirected graph of the computing power network, execution nodes for two teaching models are selected. The undirected graph of the computing power network has computing power services as nodes, network bandwidth between nodes as side lengths, and an undirected cyclic structure. The execution nodes of the two teaching models run the corresponding teaching models respectively, generating two functionally differentiated flow vector graphs; Detect and handle node overlap issues between the two flow vector graphs, assign overlapping nodes to the flow vector graph with higher fit and update it; By merging the two updated flow vector graphs, the overall flow graph of the computing power network is obtained, thus completing the networking of the dynamic computing power network.
[0131] In some embodiments, the node overlap problem is handled by: quantifying the degree of cooperative adaptation between each overlapping node and the two flow vector graphs, wherein the degree of cooperative adaptation is determined based on the link association parameters and node feature differences between the overlapping node and the non-overlapping nodes in the corresponding flow vector graph, and each overlapping node is uniquely assigned to the flow vector graph with a higher degree of cooperative adaptation.
[0132] In some embodiments, the selection of the appropriate supervisor model includes: Based on the topological characteristics of the dynamic computing network, the degree of fit between the two professor models and the dynamic computing network is calculated respectively. The professor model with a higher degree of fit is used as the computing fragmentation supervisor model, and the other professor model is used as the hardware fragmentation supervisor model.
[0133] In some embodiments, the inter-layer model collaboration mechanism includes: Clearly define the hierarchical responsibilities and powers of professors, teachers, and students, and establish a two-way collaborative link: professors issue control instructions to teachers, teachers transmit transformed instructions to students, students provide feedback on task execution status to teachers, and teachers report collaborative effects to professors.
[0134] In some embodiments, the hierarchical sample processing and distribution includes: Based on the functional differences of each level of the three-layer model, standardized labeled samples are distributed to the corresponding level model according to preset rules. The professor level guides the sample processing rules, the teacher level transmits samples and processing instructions, and the student level executes sample processing and reasoning tasks.
[0135] In some embodiments, the suitability assessment includes: after the full completion of a single round of sample training / inference tasks, two teaching models respectively conduct a quantitative assessment of the suitability of the model collaboration effect and the computing power integration strategy, and count the suitability based on the assessment results for subsequent adjustment of the master model and optimization of collaboration parameters.
[0136] In some embodiments, the dynamic adjustment includes: Based on the dispersion of the suitability assessment results, the number of student node grouping targets and the number of student layer models corresponding to each teacher layer model are dynamically optimized. At the same time, the allocation of supervisor roles for the two teaching models is adjusted to achieve adaptive optimization of the model collaboration relationship.
[0137] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0138] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 4 As shown, this disclosure also provides an electronic device 400, which includes at least one processor 401 and a memory 402 coupled to the processor 401. The memory 402 is used to store at least one processor 401 executable instructions, wherein the at least one processor 401 is used to execute the instructions to implement the steps of the method described above in this disclosure.
[0139] The processor 401 described above can also be referred to as a Central Processing Unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method described in this embodiment can be implemented by the integrated logic circuitry in the hardware of the processor 401 or by instructions in software form. The processor 401 described above can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method in conjunction with this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 402, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 401 reads information from the memory 402 and, in conjunction with its hardware, completes the steps of the method described above.
[0140] Figure 5 This is a schematic diagram of an exemplary computer system provided by an embodiment of the present disclosure. Various operations / processes according to embodiments of the present disclosure, implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, for example... Figure 5 The computer system 500 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above.
[0141] Computer system 500 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of this disclosure described and / or claimed herein.
[0142] like Figure 5As shown, the computer system 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the computer system 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0143] Multiple components in the computer system 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device capable of inputting information into the computer system 500. The input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 508 may include, but is not limited to, a hard disk and an optical disk. The communication unit 509 allows the computer system 500 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, Wi-Fi devices, WiMax devices, cellular communication devices, and / or the like.
[0144] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running analytical algorithm models, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the methods described above in the embodiments of this disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 502 and / or communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the methods described above in the embodiments of this disclosure by any other suitable means (e.g., by means of firmware).
[0145] This disclosure provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the methods described in this disclosure.
[0146] Computer-readable storage media can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0147] It should be noted that the computer-readable storage medium described in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), or any suitable combination thereof.
[0148] Embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the large model processing method described above.
[0149] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0150] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0151] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.
[0152] It should be noted that, in this document, terms such as "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0153] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for processing large models, characterized in that, The method includes: A three-layer model with interpretable logical nodes is constructed, and any target large model that can be accessed is abstracted and decomposed into computing power units that are adapted to the three-layer model. The three-layer model includes: a professor layer that realizes link architecture control and decision guidance, a teacher layer that realizes instruction conversion and transmission, and a student layer that realizes task response and implementation. Perform overall traffic graph generation and networking operations for the computing power network, and construct a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can adaptively adjust based on changes in computing power distribution. Based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer model, a suitable master model is selected, and a collaborative mechanism is established between the models at each layer, with the professor layer in charge, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer model and the dynamic computing network. Based on the aforementioned collaborative mechanism, iterative operations are performed during model training. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, the model collaboration relationship and the integration effect of large model computing power are optimized, thus completing the processing of the target large model.
2. The method according to claim 1, characterized in that, The professor layer includes two functionally differentiated professor models, each corresponding to a different execution node, used to collaboratively implement link architecture control, decision guidance, and suitability assessment functions.
3. The method according to claim 1, characterized in that, The process of abstracting and decomposing any large target model into computing power units adapted to the three-layer model includes: Based on the computing power requirements, task types, and functions of each level of the three-layer model, the target large model is decomposed into multiple computing power units that can be allocated to each level of the model. The computing power units can be adapted to the task processing requirements of the corresponding level.
4. The method according to claim 1, characterized in that, Perform overall traffic graph generation and networking operations for the computing power network, including: Based on a pre-defined undirected graph of the computing power network, execution nodes for two teaching models are selected. The undirected graph of the computing power network has computing power services as nodes, network bandwidth between nodes as side lengths, and an undirected cyclic structure. The execution nodes of the two teaching models run the corresponding teaching models respectively, generating two functionally differentiated flow vector graphs; Detect and handle node overlap issues between the two flow vector graphs, assign overlapping nodes to the flow vector graph with higher fit and update it; By merging the two updated flow vector graphs, the overall flow graph of the computing power network is obtained, thus completing the networking of the dynamic computing power network.
5. The method according to claim 4, characterized in that, The method for handling the node overlap problem is as follows: quantify the degree of cooperative adaptation between each overlapping node and the two flow vector graphs. The degree of cooperative adaptation is determined based on the link association parameters and node feature differences between the overlapping node and the non-overlapping nodes in the corresponding flow vector graph. Each overlapping node is uniquely assigned to the flow vector graph with a higher degree of cooperative adaptation.
6. The method according to claim 1, characterized in that, The selected and adapted supervisor models include: Based on the topological characteristics of the dynamic computing network, the degree of fit between the two professor models and the dynamic computing network is calculated respectively. The professor model with a higher degree of fit is used as the computing fragmentation supervisor model, and the other professor model is used as the hardware fragmentation supervisor model.
7. The method according to claim 1, characterized in that, The inter-layer model collaboration mechanism includes: Clearly define the hierarchical responsibilities and powers of professors, teachers, and students, and establish a two-way collaborative link: professors issue control instructions to teachers, teachers transmit transformed instructions to students, students provide feedback on task execution status to teachers, and teachers report collaborative effects to professors.
8. The method according to claim 1, characterized in that, The hierarchical sample processing and distribution includes: Based on the functional differences of each level of the three-layer model, standardized labeled samples are distributed to the corresponding level model according to preset rules. The professor level guides the sample processing rules, the teacher level transmits samples and processing instructions, and the student level executes sample processing and reasoning tasks.
9. The method according to claim 1, characterized in that, The suitability assessment includes: after the full completion of a single round of sample training / inference tasks, the two teaching models respectively conduct a quantitative assessment of the suitability of the model collaboration effect and the computing power integration strategy, and count the suitability based on the assessment results for subsequent adjustment of the master model and optimization of collaboration parameters.
10. The method according to claim 1, characterized in that, The dynamic adjustment includes: Based on the dispersion of the suitability assessment results, the number of student node grouping targets and the number of student layer models corresponding to each teacher layer model are dynamically optimized. At the same time, the allocation of supervisor roles for the two teaching models is adjusted to achieve adaptive optimization of the model collaboration relationship.
11. An apparatus for processing large models, characterized in that, The device includes: The building module is configured to build a three-layer model of interpretable logical nodes, which abstracts and decomposes any target large model into computing power units that adapt to the three-layer model. The three-layer model includes: a professor layer that implements link architecture control and decision guidance, a teacher layer that implements instruction conversion and transmission, and a student layer that implements task response and implementation. The execution module is configured to perform the generation of the overall traffic graph of the computing power network and the networking operation, and to build a dynamic computing power network that adapts to the integration of computing power fragments and provides computing power support for the three-layer model. The dynamic computing power network can be adaptively adjusted based on changes in computing power distribution. The filtering module is configured to filter suitable master models based on the topological characteristics of the dynamic computing network and the functional differences of the three-layer models, and establish a collaborative mechanism among the models at each layer, with the professor layer leading, the teacher layer transmitting, and the student layer executing, so as to realize the collaborative linkage between the three-layer models and the dynamic computing network. The optimization module is configured to perform iterative operations during model training based on the aforementioned collaborative mechanism. By adapting the hierarchical sample processing and distribution to the three-layer model hierarchy, evaluating the suitability of model collaboration effects, and dynamically adjusting model collaboration parameters, it optimizes the model collaboration relationship and the integration effect of large model computing power, thereby completing the processing of the target large model.
12. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-10.