Method and computing device for model optimization

By inserting synchronization operators into the multi-engine model of large models at the edge, the problems of low data processing efficiency and complex synchronization program development in parallel multi-engine execution are solved, and efficient data processing and synchronization operations are achieved, reducing the complexity of model development.

CN119645428BActive Publication Date: 2025-05-09SHENZHEN CORERAIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510168653.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-09
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

When deploying large models at the edge, parallel execution of multiple engines leads to inefficient data processing and requires complex synchronous program development to ensure the correctness of execution results, increasing the complexity of the model development process.

Method used

By loading the original model structure of the computing model, split it into a multi-engine model structure, and inserting synchronization operators into the multi-engine model structure to form a multi-engine synchronization model. The model is executed on an edge computing device, and uses synchronization operators to synchronize data and process, optimizing data processing capabilities and execution efficiency.

Benefits of technology

It realizes efficient data processing and synchronous operation of multi-engine models, reduces the complexity of the model development process, improves the efficiency of model execution, and ensures the accuracy and efficiency of calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645428B_ABST
    Figure CN119645428B_ABST
Patent Text Reader

Abstract

The present invention provides a model optimization method and computing device for a multi-engine edge computing device, the method comprising: loading the original model structure of the computing model; splitting the original model structure to obtain a multi-engine model structure; inserting a synchronization operator into the multi-engine model structure to obtain a multi-engine synchronization model; loading the multi-engine synchronization model into an edge computing device to perform model operations, the edge computing device comprising multiple computing engines, the multiple computing engines synchronously executing the multi-engine synchronization model based on the synchronization operator. According to the technical solution of the present invention, it is possible to optimize the data processing capability of the multi-engine model, improve the efficiency of model execution, and reduce the complexity of the model development process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software development, and in particular to a model optimization method and computing equipment. Background Art

[0002] With the development of deep learning, large models have made breakthroughs, and more and more large models have emerged. As lightweight large models continue to emerge, this phenomenon provides an opportunity for deploying large models on the edge. The execution of lightweight large models on the edge mainly includes direct loading and execution, splitting the model into multiple small models for execution, etc. Some edge devices with sufficient resources can be loaded and executed directly. When the resources of a single engine of an edge device are insufficient, the model needs to be split into multiple small models and loaded into different engines for execution. Model splitting includes splitting the model by segment, splitting by data, and other splitting methods.

[0003] When the model is split into multiple small models and executed in parallel by multiple engines, the model needs to be split by data. After splitting by data, the model nodes of the data that cannot be split need to perform data synchronization operations to ensure the correctness of the execution results, which requires complex synchronization program development to achieve.

[0004] To this end, a technical solution is needed that can optimize the data processing capabilities of multi-engine models, improve the efficiency of model execution, and reduce the complexity of the model development process. Summary of the invention

[0005] The present invention aims to provide a method and computing device for model optimization, which can optimize the data processing capability of a multi-engine model, improve the efficiency of model execution, and reduce the complexity of the model development process.

[0006] According to one aspect of the present invention, a model optimization method is provided for a multi-engine edge computing device, the method comprising:

[0007] Load the original model structure of the computational model;

[0008] Splitting the original model structure to obtain a multi-engine model structure;

[0009] Inserting a synchronization operator into the multi-engine model structure to obtain a multi-engine synchronization model;

[0010] The multi-engine synchronization model is loaded into an edge computing device to execute model operations, wherein the edge computing device includes multiple computing engines, and the multiple computing engines synchronously execute the multi-engine synchronization model based on the synchronization operator.

[0011] According to some embodiments, loading the original model structure of the computational model includes:

[0012] The topological information of the computing model is parsed, and the original model structure is loaded, wherein the edges in the topological information are data transfers, and the nodes are data modules.

[0013] According to some embodiments, splitting the original model structure to obtain a multi-engine model structure includes: splitting the original model structure based on data dimensions.

[0014] According to some embodiments, inserting a synchronization operator into the multi-engine model structure to obtain a multi-engine synchronization model includes: the synchronization operation is compiled and inserted into the multi-engine model structure as the synchronization operator.

[0015] According to some embodiments, the synchronization operator performs data synchronization and process synchronization.

[0016] According to some embodiments, the process synchronization includes:

[0017] When compiling the multi-engine synchronization model, the synchronization data name and the data copy name in each synchronization operator are saved in the synchronization operator attributes;

[0018] The process synchronization is completed based on the synchronization data name and the data copy name in the attributes of the synchronization operator.

[0019] According to some embodiments, the data synchronization comprises:

[0020] Writing the data information and size information of the synchronization operator into a shared data structure of all synchronization operators;

[0021] The synchronization operator copies data in the shared data structure during execution to perform data synchronization.

[0022] According to some embodiments, after all synchronously executed synchronization operators are executed and the data synchronization is completed, the multi-engine synchronization model continues to execute subsequent modules.

[0023] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method as described above is implemented.

[0024] According to another aspect of the present invention, there is provided a computing device, comprising:

[0025] processor;

[0026] A memory stores a computer program, and when the computer program is executed by the processor, the method described in any one of the above items is implemented.

[0027] According to an embodiment of the present invention, the original model structure of the computing model is first loaded, and the original model structure is split to obtain a multi-engine model structure. Through the multi-engine model structure, the position of the synchronization operator insertion can be determined. The synchronization operator is inserted into the multi-engine model structure to obtain a multi-engine synchronization model. The multi-engine synchronization model implements the synchronization operation of the model execution and processing data, optimizes the data processing capability of the multi-engine model, and improves the efficiency of model execution. After the current technology is split into multiple models, it is necessary to develop a separate synchronization program to synchronize data and processes during execution. The present invention inserts a synchronization operator into the split model. All synchronization functions are in the synchronization operator. There is no need to develop additional synchronization operation programs each time. It is only necessary to compile the model to realize data parallel synchronization processing, which improves the efficiency of model development and effectively reduces the complexity of the model development process.

[0028] According to some embodiments, the insertion of synchronization operators can ensure coordination and data consistency between different engines in a distributed environment, which is crucial for maintaining accuracy and efficiency of calculations and avoiding data inconsistency or calculation errors caused by asynchrony.

[0029] According to some embodiments, the present invention flexibly adjusts the splitting method and synchronization strategy of the model according to specific computing requirements and available resources. As computing requirements grow or hardware facilities are upgraded, the model can be relatively easily adjusted to adapt to the new environment, thereby improving the flexibility and scalability of the system.

[0030] According to some embodiments, since there is no need to customize complex data synchronization and task coordination logic for each specific application scenario, the present invention can significantly reduce development costs and shorten project cycles. By inserting synchronization operators into the split model to achieve synchronization operations between multiple engines, there is no need to develop additional complex processes or logic. Developers can more easily deploy optimized models to different computing resources, simplify the deployment process, and reduce the workload and technical difficulty when transitioning from R&D to production environments.

[0031] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below.

[0033] Figure 1 A flow chart of a method for model optimization according to an example embodiment is shown.

[0034] Figure 2 A schematic diagram illustrating an original model structure according to an example embodiment.

[0035] Figure 3 A schematic diagram illustrating a multi-engine model structure according to an example embodiment.

[0036] Figure 4 A schematic diagram illustrating a multi-engine synchronization model according to an example embodiment.

[0037] Figure 5 A block diagram of a computing device is shown according to an exemplary embodiment. DETAILED DESCRIPTION

[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present invention will be comprehensive and complete and fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.

[0039] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present invention. However, those skilled in the art will appreciate that the technical solution of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring various aspects of the present invention.

[0040] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0041] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0042] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the present inventive concept. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.

[0043] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0044] Those skilled in the art will appreciate that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present invention, and therefore cannot be used to limit the protection scope of the present invention.

[0045] In a computing environment, the failure of any node may affect the stability and correctness of the entire system. In order to use multiple engines for parallel processing, large-scale data must first be reasonably divided into multiple parts and assigned to different engines for processing. If this process is not designed properly, it may lead to data skew, that is, some engines process much more data than other engines, thus affecting overall efficiency. As the number of engines increases, the complexity of the system also increases. Configuration management, monitoring, debugging and other aspects will become more complex, placing higher demands on developers and operation and maintenance teams.

[0046] Direct loading and execution during calculation requires hardware with large computing resources to complete. There are relatively few devices with such large computing power in edge devices. When splitting a large model into multiple small models by segment, the model will be executed completely serially. Only after executing one model can the subsequent model be executed. Using multi-engine edge devices cannot efficiently utilize the hardware resources of multiple engines. When splitting into multiple small models by data, data synchronization is required at the location where the data cannot be split, which requires complex synchronization program development.

[0047] To this end, the present invention proposes a model optimization method that can optimize the data processing capability of the multi-engine model, improve the efficiency of model execution, and reduce the complexity of the model development process.

[0048] Exemplary embodiments of the present invention are described below with reference to the accompanying drawings.

[0049] Figure 1 A flow chart of a method for model optimization according to an example embodiment is shown.

[0050] See also Figure 1 , in S101, the original model structure of the calculation model is loaded.

[0051] According to some embodiments, the topological information of the computing model is parsed and the original model structure is loaded, the edges in the topological information are data transfers, and the nodes are data modules.

[0052] According to some embodiments, the original computing model is first loaded, its topological information is parsed, and the original model structure is constructed. The fragments of the original model structure are as follows: Figure 2 A schematic diagram of the original model structure is shown. Figure 2 It represents the decoding part of the large model. Module D_i represents the i-th decoding module in the original computing model, and module D_i+1 represents the i+1-th decoding module in the original computing model.

[0053] In S103, the original model structure is split to obtain a multi-engine model structure.

[0054] According to some embodiments, the original model structure is split based on data dimensions.

[0055] For the input or weight of the key operators within the model structure, the data is split in the shape dimension. The split will not affect the subsequent calculations, and will be re-aggregated at a specific node without changing the calculation results of the original calculation graph, ensuring the calculation equivalence before splitting the model structure.

[0056] For example, the original input data size for calculation is 1x32x512x512, which is split into two 1x32x512x256 data according to the data dimension, and put into two engines for calculation respectively, and then the results are merged during the synchronization process.

[0057] According to some embodiments, the data splitting model is a multi-engine model, and a data module in the original computing model is split into multiple data modules for multi-engine model computing. Figure 3 A schematic diagram showing a calculation model split into two engine models is shown, and the modules in the calculation model are also split into two modules. The two split modules need to be synchronized after execution. For example, module D_i is split into module D_i' and module D_i''. After the two split modules are executed, synchronization is required to continue to calculate and execute subsequent modules.

[0058] According to some embodiments, Figure 3 The original computing model is compiled for the first time, split into a multi-engine model, and the data nodes that need to be synchronized are determined. The figure shows which data nodes need to be synchronized.

[0059] In S105, a synchronization operator is inserted into the multi-engine model structure to obtain a multi-engine synchronization model.

[0060] According to some embodiments, the synchronization operation is inserted into the multi-engine model structure as a model node, and the synchronization operation node becomes the synchronization operator after being compiled. The synchronization operator performs data synchronization and process synchronization. After all synchronization operators executed synchronously are completed and the data synchronization is completed, the multi-engine synchronization model continues to execute subsequent modules.

[0061] According to some embodiments, a synchronization operator is developed based on a synchronization operator process and provided to a multi-engine synchronization model for use in performing calculations.

[0062] According to some embodiments, when compiling the multi-engine synchronization model, the synchronization data name and the data copy name in each synchronization operator are saved to the synchronization operator attributes, and the process synchronization is completed based on the synchronization data name and the data copy name in the synchronization operator attributes.

[0063] According to some embodiments, the data information and size information of the synchronization operator are written into a shared data structure of all synchronization operators, and the synchronization operator copies data in the shared data structure during execution to perform data synchronization.

[0064] In the traditional multi-engine model execution process, users need to obtain information in the model and then develop synchronous operation programs to implement model calculations. The development process is highly complex and the development cost is too high. Figure 4 The present invention compiles the model based on the multi-engine model, converts the synchronization operation originally required by the user into a model node or module inserted into the model structure, and provides a synchronization operator to realize the multi-engine synchronization model to perform large-scale data calculations.

[0065] The synchronization operator involves process synchronization and data synchronization between multiple engines. When compiling the model, the name of the data that needs to be synchronized and the name of the data copy in each synchronization operator are saved in the synchronization operator attributes. When the synchronization operator is executed, process synchronization and data synchronization are performed based on the name of the synchronized data and the name of the data to be copied.

[0066] According to some embodiments, Figure 4 After the multi-engine model is compiled for the second time, the nodes that need to be synchronized are inserted into the multi-engine model as synchronization operators, and the node positions of the synchronization operators in the model are shown.

[0067] The process synchronization is completed based on the name list of the data that needs to be synchronized in the synchronization operator attributes, and the data offset and data size corresponding to the data in the name list are copied to complete the data synchronization. Process synchronization means that all batches of synchronization operators that need to be synchronized are executed and data synchronization is completed before the subsequent modules can be executed. Each synchronization operator will provide its own data and size information and write it into the data structure shared by all synchronization operators. When other synchronization operators are executed, they copy the data in the shared data structure for data synchronization.

[0068] In S107, the multi-engine synchronization model is loaded into an edge computing device to execute model operations, wherein the edge computing device includes multiple computing engines, and the multiple computing engines synchronously execute the multi-engine synchronization model based on the synchronization operator.

[0069] According to some embodiments, the multi-engine synchronization model is executed, and the execution result of the multi-engine synchronization model is compared with the execution result of the calculation model. If they are the same, it is confirmed that the model optimization method is correct.

[0070] Load the multi-engine synchronization model with the synchronization operator inserted, and execute the model to verify the results. Load the multi-engine synchronization model with the synchronization operator inserted, and execute all multi-engine models in the same process through the synchronization operator to obtain the final result. Compare the final result with the execution result of the original model in the same process. If the results are consistent, it is confirmed that the optimization method is correct, and it is verified that the synchronization operator in the multi-engine model with the synchronization operator inserted in the present invention is inserted correctly and the original model is split correctly.

[0071] Figure 5 A block diagram of a computing device is shown according to an exemplary embodiment.

[0072] like Figure 5 As shown, computing device 30 includes processor 12 and memory 14. Computing device 30 may also include bus 22, network interface 16, and I / O interface 18. Processor 12, memory 14, network interface 16, and I / O interface 18 may communicate with each other via bus 22.

[0073] The processor 12 may include one or more general-purpose CPUs (Central Processing Units, processors), microprocessors, or application-specific integrated circuits, etc., for executing relevant program instructions. According to some embodiments, the computing device 30 may also include a high-performance graphics card (GPU) 20 for accelerating the processor 12.

[0074] The memory 14 may include a machine system readable medium in the form of a volatile memory, such as a random access memory (RAM), a read-only memory (ROM) and / or a cache memory. The memory 14 is used to store one or more programs including instructions and data. The processor 12 may read the instructions stored in the memory 14 to execute the above-mentioned method according to the embodiment of the present invention.

[0075] The computing device 30 may also communicate with one or more networks via the network interface 16. The network interface 16 may be a wireless network interface.

[0076] The bus 22 may include an address bus, a data bus, a control bus, etc. The bus 22 provides a path for exchanging information between components.

[0077] It should be noted that, in the specific implementation process, the computing device 30 may also include other components necessary for normal operation. In addition, those skilled in the art may understand that the above device may only include components necessary for implementing the embodiments of this specification, and need not include all components shown in the figure.

[0078] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the program is executed by a processor. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a micro drive, and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), a network storage device, a cloud storage device, or any type of medium or device suitable for storing instructions and / or data.

[0079] An embodiment of the present invention further provides a computer program product, which includes a computer program. The computer program is operable to enable a computer to execute part or all of the steps of any one of the methods described in the above method embodiments.

[0080] Those skilled in the art can clearly understand that the technical solution of the present invention can be implemented with the help of software and / or hardware. "Unit" and "module" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware can be, for example, a field programmable gate array, an integrated circuit, etc.

[0081] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0082] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0083] In the several embodiments provided by the present invention, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0084] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0086] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention.

[0087] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0088] The exemplary embodiments of the present invention are specifically shown and described above. It should be understood that the present invention is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present invention is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended clauses.

Claims

1. A model optimization method for a multi-engine edge computing device, characterized in that: The method comprises: Load the original model structure of the computational model; The original model structure is split according to data to obtain a multi-engine model structure, and the synchronization operation is inserted into the multi-engine model structure as a model node, and the node of the synchronization operation is compiled to become a synchronization operator; Inserting the synchronization operator into the multi-engine model structure to obtain a multi-engine synchronization model; Loading the multi-engine synchronization model to an edge computing device to perform model operations, the edge computing device comprising a plurality of computing engines, the plurality of computing engines synchronously executing the multi-engine synchronization model based on the synchronization operator; The synchronization operator performs data synchronization and process synchronization. After all synchronization operators that are executed synchronously are completed and the data synchronization is completed, the multi-engine synchronization model continues to execute subsequent modules.

2. The method according to claim 1, characterized in that Load the original model structure of the calculation model, including: The topological information of the computing model is parsed, and the original model structure is loaded, wherein the edges in the topological information are data transfers, and the nodes are data modules.

3. The method according to claim 1, characterized in that: The process synchronization includes: When compiling the multi-engine synchronization model, the synchronization data name and the data copy name in each synchronization operator are saved in the synchronization operator attributes; The process synchronization is completed based on the synchronization data name and the data copy name in the attributes of the synchronization operator.

4. The method according to claim 1, characterized in that: The data synchronization includes: Writing the data information and size information of the synchronization operator into a shared data structure of all synchronization operators; The synchronization operator copies data in the shared data structure during execution to perform data synchronization.

5. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 4 when being executed by a processor.

6. A computing device, characterized in that include: processor; A memory storing a computer program, wherein when the computer program is executed by the processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method and device for predicting training time consumption in heterogeneous computing power based on CXL

    CN119204361A