Model Data Processing Method, Apparatus, Computer Device, and Storage Medium
By decoupling and discretizing the pre-trained model of the state space model, a second operator suitable for other hardware processing units is generated, which solves the problem of slow data processing speed in the prior art and achieves a more efficient data processing effect.
Patent Information
- Application Number
- CN202510309307.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The current technology has low data processing speed and efficiency under the state space model structure, and it is impossible to significantly improve the data processing speed through technologies such as operator fusion, quantization and low-rank decomposition.
By decoupling the pretrained model under the state space model structure, a second operator is generated for other hardware processing units, and the data to be processed is discretized to reduce complex calculations.
The data processing speed and efficiency of the state space model are significantly improved, and more efficient data processing is achieved by using the computing power of other hardware processing units to accelerate computing and discrete processing.
Smart Images

Figure CN119830979B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a model data processing method, apparatus, computer device, and storage medium. Background Art
[0002] In the field of artificial intelligence, data inference of a model is a key link, and the data processing speed of the model will directly affect the data processing efficiency of the model. During the data processing of the model, various types of operators are often involved, and the data is processed through the operators. Usually, the operators perform continuous calculations during the data processing, which results in low computational efficiency of the hardware architecture for running the model.
[0003] In related technologies, techniques such as operator fusion, quantization, model pruning, and low-rank decomposition are adopted to improve the data processing efficiency of the model. However, the data processing acceleration effects of these technologies are limited during the data processing of models with some specific structures. For example, the data processing acceleration effect for the visual state space model under the state space model structure (SSM) is limited.
[0004] Therefore, for models under the state space model structure, the acceleration effect of data processing for such models in related technologies is low, and the data processing speed of such models cannot be significantly improved, resulting in slow and low-efficiency data processing of such models. Summary of the Invention
[0005] Embodiments of this application provide a model data processing method, apparatus, computer device, and storage medium, which can significantly improve the data processing speed and processing efficiency of models under the state space model structure.
[0006] To achieve the above objective, on the one hand, an embodiment of this application provides a model data processing method, including:
[0007] Obtain a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under the state space model structure;
[0008] Determine a first operator in the pre-trained model that needs to be accelerated in data processing according to the model data;
[0009] Decouple the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator;
[0010] Obtain the data to be processed corresponding to the second operator, and determine the discretization parameters corresponding to the data to be processed;
[0011] Perform data processing on the data to be processed and the discretization parameters according to the second operator to obtain a data processing result.
[0012] To achieve the above object, an embodiment of the present application provides a model data processing device on the one hand, including:
[0013] A first acquisition module, configured to acquire a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under a state space model structure;
[0014] A determination module, configured to determine a first operator that needs to be accelerated in data processing in the pre-trained model according to the model data;
[0015] A decoupling module, configured to decouple the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator;
[0016] A second acquisition module, configured to acquire the data to be processed corresponding to the second operator, and determine the discretization parameters corresponding to the data to be processed;
[0017] A processing module, configured to perform data processing on the data to be processed and the discretization parameters according to the second operator to obtain a data processing result.
[0018] In some embodiments, the decoupling module is configured to:
[0019] Determine the operation logic corresponding to each sub-operator in the first operator;
[0020] Determine the type of hardware processing unit corresponding to each sub-operator according to the operation logic;
[0021] Decouple each sub-operator according to the type of hardware processing unit to obtain a second operator.
[0022] In some embodiments, the decoupling module is configured to:
[0023] Determine the operation logic corresponding to each sub-operator in the first operator;
[0024] Modify each of the sub-operators according to a pre-designed calculation library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator;
[0025] Generate a second operator corresponding to the first operator according to the target sub-operator corresponding to each sub-operator.
[0026] In some embodiments, the processing module is configured to:
[0027] Determine a state transition matrix, an input control matrix, and an output control matrix corresponding to the second operator;
[0028] Determine a discrete step matrix according to the discretization parameter, multiply the discrete step matrix by the state transition matrix, and perform exponentiation processing on each element in the multiplied matrix to obtain a discrete state transition matrix;
[0029] Determine an input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a discrete input control matrix;
[0030] Obtain a state matrix corresponding to the previous moment, multiply the state matrix by the discrete state transition matrix and then add the discrete input control matrix to obtain a target state matrix corresponding to the current moment;
[0031] Multiply the target state matrix by the output control matrix to obtain an output matrix corresponding to the current moment;
[0032] Determine a data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment.
[0033] In some embodiments, the processing module is configured to:
[0034] Determine an output matrix corresponding to an input matrix at other moments in the data to be processed;
[0035] Perform matrix stacking processing on the output matrix corresponding to the current moment and the output matrix corresponding to the other moments to obtain a first target output matrix corresponding to the data to be processed, and the first target output matrix is the data processing result corresponding to the data to be processed.
[0036] In some embodiments, the processing module is configured to:
[0037] Perform data format conversion on the second operator and other operators corresponding to the pre-trained model to obtain a target pre-trained model in a target data format;
[0038] Determine a third operator corresponding to the second operator in the target pre-trained model;
[0039] Perform data processing on the data to be processed and the discretization parameters according to the third operator to obtain a data processing result.
[0040] In some embodiments, the processing module is configured to:
[0041] Determine the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator;
[0042] Perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain a first processing matrix, and perform one-dimensional convolution processing on the first processing matrix and the state transition matrix to obtain a target state matrix;
[0043] Multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a first target input matrix;
[0044] Perform block accumulation processing on the first target input matrix to obtain a second target input matrix, and determine the output matrix to be processed according to the second target input matrix, the target state matrix, and a preset state space matrix with all elements being zero;
[0045] Perform one-dimensional convolution processing on the output matrix to be processed and the output control matrix to obtain a second target output matrix, and the second target output matrix is the data processing result corresponding to the data to be processed.
[0046] In some embodiments, the processing module is configured to:
[0047] Perform block processing on the input matrix corresponding to the data to be processed to obtain a plurality of sub-matrices;
[0048] Determine the cumulative sum matrix corresponding to each sub-matrix according to the time step of each sub-matrix, and perform matrix stacking processing on the cumulative sum matrices corresponding to each sub-matrix to obtain a first processing matrix.
[0049] In some embodiments, the processing module is configured to:
[0050] Determine the sub-matrix corresponding to each time step in the first target input matrix;
[0051] Accumulate the sub-matrices corresponding to each time step to obtain a second target input matrix.
[0052] To achieve the above object, an aspect of the embodiments of the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the model data processing method provided by the embodiments of the present application.
[0053] To achieve the above object, on the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model data processing method provided by the embodiment of the present application is implemented.
[0054] In the embodiment of the present application, by obtaining a pre-trained model and model data corresponding to the pre-trained model, the pre-trained model is a model under a state space model structure; determining a first operator in the pre-trained model that needs to be accelerated in data processing according to the model data; decoupling the first operator to obtain a second operator corresponding to the first operator, and the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; obtaining the data to be processed corresponding to the second operator, and determining the discretization parameter corresponding to the data to be processed; performing data processing on the data to be processed and the discretization parameter according to the second operator to obtain a data processing result.
[0055] In this way, by determining the first operator that needs to be accelerated in data processing from the model data of the pre-trained model under the state space model structure, and then decoupling the first operator, the second operator can be obtained. The second operator can be applied to other hardware processing units relative to the first operator. In this way, the computing power of other hardware processing units can be utilized to accelerate the operation of the second operator. Moreover, the data to be processed that needs to be processed by the second operator is discretized using the discretization parameter. In this way, the second operator avoids directly processing the data to be processed as a whole, but processes the discretized data. This can reduce the complex calculation of the second operator on the overall data to be processed, thereby improving the processing speed of the second operator for the data to be processed. Therefore, compared with the limited data acceleration processing effect of the model acceleration technology in the related art for the model under the state space model structure, in the present application, the first operator can be decoupled to obtain the second operator, and then the second operator can be loaded by other hardware units, and the data to be processed of the second operator is discretized, which can significantly improve the data processing speed of the second operator, thereby significantly improving the data processing speed of the pre-trained model of the entire state space model. Therefore, the data processing acceleration effect on the pre-trained model of the state space model is more obvious than that in the related art.
[0056] Other features and advantages of the present application will be described in the following description, and, in part, will be obvious from the description, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0058] Figure 1 It is a schematic diagram of the system framework corresponding to the model data processing method provided by the embodiment of the present application;
[0059] Figure 2 It is a schematic flowchart of the model data processing method provided by the embodiment of the present application;
[0060] Figure 3 It is a schematic flowchart of the steps included in step 230 provided by the embodiment of the present application;
[0061] Figure 4 It is another schematic flowchart of the steps included in step 230 provided by the embodiment of the present application;
[0062] Figure 5 It is a comparison chart of the number of operators corresponding to the pre-trained model and the target pre-trained model provided by the embodiment of the present application;
[0063] Figure 6 It is another schematic flowchart of the model data processing method provided by the embodiment of the present application;
[0064] Figure 7 It is a schematic diagram of the structure of the model data processing device provided by the embodiment of the present application;
[0065] Figure 8 It is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0066] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by the present application.
[0067] It should be noted that in the specific implementation manners of the present application, regarding data related to the model, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards.
[0068] It should be noted that in some processes described in the specification, claims, and the above-mentioned drawings, multiple steps appear in a specific order. However, it should be clearly understood that these steps can be executed not in the order in which they appear herein or in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this article are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0069] The model data processing method provided by the embodiments of this application relates to the field of artificial intelligence technology. The model data processing method provided by the embodiments of this application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the model data processing method, etc., but is not limited to the above forms.
[0070] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0071] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations:
[0072] State Space Model (SSM): A state space model is a mathematical model of a dynamic system that describes the dynamic behavior of a system through state variables. A state space model mainly consists of two equations: the state equation and the observation equation.
[0073] Visual State Space Model: It is the application of the state space model in the field of vision. It regards the visual system (such as tasks like object recognition, tracking, and scene understanding in computer vision) as a dynamic system, which also includes a state equation and an observation equation.
[0074] Vmamba model: It is a model corresponding to the state space model structure. Mainly inspired by the ability of the state space model in long-sequence modeling, it introduces related concepts of the state space model into the field of vision. Its core contains a series of Visual State Space (VSS) blocks, which process visual data using unique calculation methods.
[0075] Operator: In the context of a model, an operator is an abstract representation of an operation on data. It can be regarded as a function or a transformation rule that acts on the input data to produce corresponding output. An operator receives data in a specific format (such as vectors, tensors, etc.) as input and, after a series of predefined calculation steps, outputs the processed data.
[0076] Discretization: Discretization is a process of converting continuous objects (such as functions, spaces, time, data, etc.) into discrete forms. It has wide applications in multiple fields such as mathematics and computer science. Its core idea is to divide or approximate continuous quantities into discrete units through specific methods, enabling complex continuous problems to be represented and calculated in a more easily processed discrete form.
[0077] ONNX (Open Neural Network Exchange): ONNX is an open format for representing deep learning models. Its main purpose is to provide a unified model representation standard, enabling convenient model conversion and interoperability between different deep learning frameworks (such as PyTorch, TensorFlow, etc.). For example, a model trained in PyTorch can be converted to the ONNX format and then deployed and inferred on other inference engines that support ONNX (such as ONNX Runtime).
[0078] Regarding other related terms that need to be explained, they will be introduced together in the following text.
[0079] First, introduce the technical problems existing in the related technologies:
[0080] In the field of artificial intelligence, data inference of a model is a key link, and the data processing speed of the model will directly affect the data processing efficiency of the model. During the data processing of the model, various types of operators are often involved, and the data processing is realized through the operators. Usually, the operators perform continuous calculations during the data processing, which results in low computational efficiency of the hardware architecture for running the model.
[0081] In related technologies, techniques such as operator fusion, quantization, model pruning, and low-rank decomposition are adopted to improve the data processing efficiency of the model. However, the data processing acceleration effects of these techniques during the data processing of models with some specific structures are limited. For example, the data processing acceleration effect for the visual state space model under the state space model structure (State Space Model, SSM) is limited.
[0082] Therefore, for the models under the state space model structure, the acceleration effect of data processing for such models in related technologies is low, and the data processing speed of such models cannot be significantly improved, resulting in slow and inefficient data processing of such models.
[0083] To solve the above technical problems, embodiments of the present application provide a model data processing method, apparatus, computer device, and storage medium.
[0084] Among them, a first operator that requires data processing acceleration is determined from the model data of a pre-trained model under the state space model structure, and then the first operator is decoupled to obtain a second operator. The second operator can be applied to other hardware processing units relative to the first operator. In this way, the computing power of other hardware processing units can be utilized to accelerate the operation of the second operator. Moreover, the data to be processed by the second operator is discretized using discretization parameters. This enables the second operator to avoid directly processing the data to be processed as a whole, but rather to process the discretized data, reducing the complex calculations of the second operator on the overall data to be processed, thereby improving the processing speed of the second operator for the data to be processed. Therefore, compared with the limited data acceleration effect of the model acceleration technology in the related art on the model data under the state space model structure, in this application, the first operator can be decoupled to obtain the second operator, and then the second operator can be loaded by other hardware units, and the data to be processed by the second operator is discretized, which can significantly improve the data processing speed of the second operator, and thus significantly improve the data processing speed of the pre-trained model of the entire state space model. Therefore, the data processing acceleration effect on the pre-trained model of the state space model is more obvious than that in the related art.
[0085] The method, device, computer device, and storage medium for model data processing will be introduced in detail later.
[0086] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the system framework corresponding to the model data processing method provided by the embodiments of the present application. The model data processing method provided by the embodiments of the present application can be applied to this system framework.
[0087] It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0088] The terminal 140 or the server 110 can be a device that executes the model data processing method.
[0089] The terminal 140 includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc. The embodiments of the present application can be applied to various scenarios, including but not limited to image classification, object detection, semantic segmentation, etc. Additionally, it can be a single device or a collection of multiple devices combined. For example, multiple desktop computers are interconnected through a local area network and share a display, etc. to work collaboratively, jointly constituting a terminal 140. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0090] Server 110 refers to a computer system that can provide certain services to terminal 140. Compared with ordinary terminal 140, server 110 has higher requirements in terms of stability, security, performance, etc. Server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0091] Gateway 120 is also called an internetwork connector and protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems that use different communication protocols, data formats or languages, and even have completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by terminal 140 to server 110 needs to be sent to the corresponding server 110 through gateway 120. The message sent by server 110 to terminal 140 also needs to be sent to the corresponding terminal 140 through gateway 120.
[0092] The model data processing method in the embodiments of this application can be applied to various scenarios, such as image classification, object detection, semantic segmentation and other scenarios. The scenarios to which the model data processing method in this application is applied are not limited herein.
[0093] Please refer to Figure 2 , Figure 2 which is a schematic flow diagram of the model data processing method provided by the embodiments of this application.
[0094] The model data processing method provided by the embodiments of this application may include the following steps:
[0095] Step 210, obtain a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under the state space model structure;
[0096] Step 220, determine a first operator that needs to be accelerated in data processing in the pre-trained model according to the model data;
[0097] Step 230, perform decoupling processing on the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator;
[0098] Step 240, obtain the data to be processed corresponding to the second operator and determine the discretization parameters corresponding to the data to be processed;
[0099] Step 250, perform data processing on the data to be processed and the discretization parameters according to the second operator to obtain a data processing result.
[0100] The following will describe steps 210 to 250 in detail.
[0101] In step 210, a pre-trained model and model data corresponding to the pre-trained model are obtained. The pre-trained model is a model under the state space model structure.
[0102] In this application, the pre-trained model can be a model under the state space model structure. The pre-trained model contains model data, and the model data can include data of multiple operators corresponding to the pre-trained model. The corresponding multiple operators in the pre-trained model can be determined through the model data.
[0103] The pre-trained model is mainly trained on the architecture of a Graphics Processing Unit (GPU), such as the VMamba model. If it is desired to migrate to other architectures for inference, some operators in the pre-trained model need to be decoupled and transformed so that these operators are suitable for being loaded and run on other hardware processing units, such as being loaded and run on a Central Processing Unit (CPU).
[0104] In step 220, a first operator that needs to be accelerated in data processing in the pre-trained model is determined according to the model data.
[0105] Multiple operators corresponding to the pre-trained model can be determined according to the model data. Since the pre-trained model is trained on the architecture of a graphics processor, most of these operators are suitable for being loaded and run on a graphics processor. However, when the graphics processor is busy with tasks and has a large load of computational work, the processing speed of the data to be processed input into the pre-trained model will be slow.
[0106] Among the multiple operators of the pre-trained model, some operators are constructed based on functions of the PyTorch framework and thus support being loaded and run on different hardware processing units. PyTorch has a set of abstract device management mechanisms. It abstractly encapsulates different hardware processing units (such as CPUs and GPUs). At the bottom layer, PyTorch defines the torch.device class, and this class can be used to specify the device where a tensor is located. For example, the tensor can be specified to operate on the CPU or GPU respectively through torch.device("cpu") or torch.device("cuda:0") (assuming there is a GPU device with the number 0). This kind of abstraction enables the operator to be independent of the specific hardware device. As long as the device supports the device abstraction layer interface of PyTorch, the operator can run on it.
[0107] However, in a pre-trained model, some custom operators or operators not built based on the PyTorch framework often rely on a graphics processing unit for processing and cannot be directly loaded and processed by other hardware processing units. These operators are determined from the model data and identified as the first operators.
[0108] To improve the data processing speed of the first operators during the calculation process, in this application, the first operators that require data processing acceleration in the pre-trained model can be determined according to the model data, and subsequent processing can be performed on the first operators, so as to achieve migrating at least some sub-operators in the first operators to other hardware processing units for loading, such as migrating to a central processing unit for loading and processing, to relieve the computing pressure on the graphics processing unit and enable the graphics processing unit and the central processing unit to perform parallel processing on the data that needs to be processed by the pre-trained model, thereby improving the data processing speed and efficiency of the pre-trained model.
[0109] For another example, without changing the calculation purpose and calculation method of the first operator, it is possible to reduce the number of sub-operators in the first operator, thereby obtaining a second operator. The second operator has fewer sub-operators and can significantly reduce the data that needs to be processed in the intermediate steps during the data processing. These will be described later.
[0110] In step 230, decouple the first operator to obtain a second operator corresponding to the first operator. The hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator.
[0111] Among them, the first operator can be the CrossScanTriton operator. The CrossScanTriton operator is an operator used for scanning operations across different dimensions. When processing images or other multi-dimensional data, traditional linear scanning may not be sufficient to capture enough context information. CrossScanTriton uses a Triton accelerator to achieve efficient cross-dimensional scanning, which can scan data from different directions or paths to more comprehensively understand and process the spatial relationships of the data.
[0112] Usage of the CrossScanTriton operator in the pre-trained model: The CrossScanTriton operator is used in the pre-trained model to achieve two-way scanning or multi-directional scanning, helping the pre-trained model obtain more extensive context information when processing visual data and enhancing the model's understanding of the relationships between different positions in the image.
[0113] The CrossScanTriton operator runs based on the Triton accelerator, which is implemented based on a graphics processing unit. If we want to migrate the CrossScanTriton operator to other hardware processing units for computing, we need to decouple the CrossScanTriton operator to obtain the corresponding decoupled second operator. The number of second operators can be one, two, or more, so that the second operator can be applied to other hardware processing units, such as a central processing unit.
[0114] The first operator can be the CrossMergeTriton operator, which is used to merge or fuse features from different scan directions or different levels. Leveraging the parallel computing power of the Triton accelerator, this operator can efficiently process and merge data, ensuring that information from different sources can be effectively integrated to form a more representative feature representation.
[0115] In visual tasks, a pre-trained model needs to integrate information from different scan paths or features at different levels. The CrossMergeTriton operator can help the pre-trained model perform this complex feature fusion while maintaining efficiency, thereby improving the model's performance. The CrossMergeTriton operator is suitable for loading and running on a graphics processing unit and needs to be decoupled to obtain the corresponding decoupled second operator. The number of second operators can be one, two, or more, so that the second operator can be applied to other hardware processing units, such as a central processing unit.
[0116] The first operator can be the SelectiveScanOflex operator. SelectiveScanOflex is an implementation of selective scanning that can dynamically select the scan path or method according to the characteristics of the input data or the requirements of the model. For example, it can select to scan from left to right, from right to left, or bidirectionally, or perform selective scanning in different dimensions. Different from the traditional fixed scan path, this operator can select different scan methods according to the characteristics of the input data or the requirements of the model to better capture important information in the data.
[0117] In a pre-trained model, such as the VMamba model, the SelectiveScanOflex operator can be used to enhance the flexibility of the model, enabling it to adjust the scanning strategy according to different input images or task requirements, thereby improving the generalization ability and efficiency of the model in different scenarios. The SelectiveScanOflex operator is suitable for loading and running on a graphics processing unit and needs to be decoupled to obtain the corresponding decoupled second operator. The number of second operators can be one, two, or more, so that the second operator can be applied to other hardware processing units, such as a central processing unit.
[0118] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the steps included in step 230 provided by an embodiment of the present application. In some embodiments, decoupling the first operator to obtain the second operator corresponding to the first operator includes the following steps:
[0119] Step 301, determine the operation logic corresponding to each sub-operator in the first operator;
[0120] Step 302, determine the type of hardware processing unit corresponding to each sub-operator according to the operation logic;
[0121] Step 303, decouple each sub-operator according to the type of hardware processing unit to obtain the second operator.
[0122] The following will describe steps 301 to 303 in detail.
[0123] In step 301, determine the operation logic corresponding to each sub-operator in the first operator.
[0124] Among them, the first operator contains multiple sub-operators, and the operation logic of each sub-operator can be analyzed. Among them, some operation logics are based on a graphics processing unit, such as parallel computing logic. Some operation logics are based on a central processing unit, such as precise computing logic.
[0125] In step 302, determine the type of hardware processing unit corresponding to each sub-operator according to the operation logic.
[0126] The type of hardware processing unit corresponding to each sub-operator can be determined according to the operation logic, that is, the hardware processing unit corresponding to some sub-operators is a graphics processing unit, and the hardware processing unit corresponding to some sub-operators is a central processing unit.
[0127] In step 303, decouple each sub-operator according to the type of hardware processing unit to obtain the second operator.
[0128] Then, each sub-operator is decoupled according to the type of hardware processing unit to obtain a second operator. Among them, the loading of some sub-operators must depend on the graphics processing unit, and these types of operators maintain their original running and loading methods. While the running and loading of some sub-operators can be performed through the graphics processing unit or the central processing unit, and they are determined as the second operator. Subsequently, these two different types of sub-operators can be loaded onto the graphics processing unit and the central processing unit respectively, thereby realizing the transfer of all the computing tasks of the original first operator on the graphics processing unit to the central processing unit, thus improving the data processing efficiency.
[0129] Please refer to Figure 4 , Figure 4 which is another schematic flowchart of the steps included in step 230 provided by the embodiments of the present application. In some embodiments, decoupling the first operator to obtain the second operator corresponding to the first operator includes:
[0130] Step 401, determining the operation logic corresponding to each sub-operator in the first operator;
[0131] Step 402, modifying each sub-operator according to the pre-designed computing library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator;
[0132] Step 403, generating a second operator corresponding to the first operator according to the target sub-operator corresponding to each sub-operator.
[0133] The following will explain steps 401 to 403 in detail.
[0134] In step 401, the operation logic corresponding to each sub-operator in the first operator is determined.
[0135] Among them, the operation logic corresponding to each sub-operator in the first operator can be determined. Specifically, analyze the input, output, and specific functions of each sub-operator in the first operator, such as whether it is used for matrix multiplication, convolution operation, or other specific mathematical operations, so as to determine the operation logic of each sub-operator.
[0136] In step 402, each sub-operator is modified according to the pre-designed computing library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator.
[0137] Then, each sub-operator is modified according to the pre-designed computing library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator. The pre-designed computing library is specifically a pre-designed computing library defined under the PyTorch framework, which includes various basic mathematical operators, such as operators for basic arithmetic operations including addition, subtraction, multiplication, and division.
[0138] After determining the operation logic of each sub-operator, a target sub-operator with the same function and role as the original sub-operator can be constructed through a pre-designed computing library. The target sub-operator has the same data input, data processing logic, role, and function as the original sub-operator, except that it is defined through the PyTorch framework. This is to achieve the conversion of atomic operators and make the target sub-operator more compatible with other hardware processing units.
[0139] In step 403, a second operator corresponding to the first operator is generated according to the target sub-operator corresponding to each sub-operator.
[0140] Then, a second operator corresponding to the first operator is generated according to the target sub-operator corresponding to each sub-operator. For example, the target sub-operators corresponding to each sub-operator can be integrated according to the overall operation logic of the first operator. For example, the inputs and outputs of different target sub-operators are connected according to the overall operation logic of the first operator, thereby generating a second operator.
[0141] Among them, the hardware processing units of the first operator and the second operator can be different. For example, the hardware processing unit of the first operator is a graphics processor, while the hardware processing unit of the second operator can be a central processing unit.
[0142] During the process of data processing by the pre-trained model, the second operator corresponding to the first operator of the pre-trained model can be handed over to the central processing unit for operation. This can share the computing tasks of the graphics processor, improve the computing efficiency of the graphics processor, and at the same time, the central processing unit and the graphics processor process the data to be processed corresponding to the pre-trained model simultaneously, which can improve the data processing speed and processing efficiency.
[0143] In step 240, the data to be processed corresponding to the second operator is obtained, and the discretization parameters corresponding to the data to be processed are determined.
[0144] The data to be processed corresponding to the second operator can be continuous data, such as a continuous matrix corresponding to continuous image data. The discretization parameters corresponding to the data to be processed can be determined, and the discretization parameters can discretize the continuous data to be processed.
[0145] Discretization usually involves converting continuous data or operations into a discrete form. In a deep learning model, this means mapping continuous weights, activation values, or other parameters from the continuous real number space to a discrete space. For example, converting floating-point weights to fixed-point numbers, or approximating continuous activation functions (such as Sigmoid or ReLU) as piecewise linear functions or step functions. This conversion can simplify some originally complex computational operations, thereby reducing the number of required operators. For example, for the discretization of certain activation functions, if the continuous Sigmoid function is approximated as a step function, the complex mathematical operations originally used to calculate the Sigmoid function can be simplified to simple comparison and assignment operations. It may not be necessary to use a dedicated activation function operator, or a simpler logical judgment operator can be used instead, thus reducing the operators in the model.
[0146] In step 250, the data to be processed and the discretization parameters are processed according to the second operator to obtain a data processing result.
[0147] Among them, the data to be processed and the discretization parameters can be input into the second operator for data processing to obtain a data processing result. Or, the pre-trained model can be converted into the ONNX data format to obtain the converted target pre-trained model. Then, the third operator corresponding to the second operator is determined in the target pre-trained model, and the data to be processed and the discretization parameters are processed according to the third operator to obtain a data processing result.
[0148] In some embodiments, processing the data to be processed and the discretization parameters according to the second operator to obtain a data processing result includes:
[0149] (1.1) Determine the state transition matrix, input control matrix, and output control matrix corresponding to the second operator;
[0150] (1.2) Determine the discrete step matrix according to the discretization parameters, multiply the discrete step matrix by the state transition matrix, and perform exponentiation on each element in the multiplied matrix to obtain the discrete state transition matrix;
[0151] (1.3) Determine the input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain the discrete input control matrix;
[0152] (1.4) Obtain the state matrix corresponding to the previous moment, multiply the state matrix by the discrete state transition matrix and add it to the discrete input control matrix to obtain the target state matrix corresponding to the current moment;
[0153] (1.5) Multiply the target state matrix by the output control matrix to obtain the output matrix corresponding to the current moment;
[0154] (1.6) Determine the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment.
[0155] Among them, determine the state transition matrix, input control matrix, and output control matrix corresponding to the second operator. The state transition matrix corresponding to the second operator is denoted as A(d_in, n), the input control matrix is denoted as B(b, l, n), and the output control matrix is denoted as C(b, l, n). These matrices are operator units inside the second operator.
[0156] The original data corresponding to the data to be processed can be image data. b represents the size of the bachsize of the image, l represents the number of image blocks, d_in represents the feature dimension of each matrix, and n is a hyperparameter.
[0157] The state transition matrix describes how the state changes over time. The input control matrix describes how external inputs affect the change of the state. The output control matrix describes how the state matrix processed inside the second operator is mapped to the output space.
[0158] Then, determine the discrete step matrix according to the discretization parameter, which can be denoted as delta(b, l, d_in). The discrete step matrix is used to discretize the data to be processed. Then multiply the discrete step matrix by the state transition matrix and exponentiate each element in the multiplied matrix to obtain the discrete state transition matrix, which can be denoted as deltaA(b, l, d_in, n).
[0159] Next, determine the input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain the discrete input control matrix. The input matrix is denoted as u(b, l, d_in). Then multiply the dimensions of the input matrix u(b, l, d_in), the input control matrix B(b, l, n), and the discrete state transition matrix deltaA(b, l, d_in, n) to obtain the discrete input control matrix deltaB_u(b, l, d_in, n).
[0160] By obtaining the state matrix corresponding to the previous moment, which is the matrix generated when the second operator processes the input matrix at the previous moment of the data to be processed and can be expressed as X(t-1), where t-1 represents the previous moment. Then, multiply the state matrix by the discrete state transition matrix and add the discrete input control matrix to obtain the target state matrix corresponding to the current moment, which can be specifically expressed as: X(t ) = deltaAx(t-1) + deltaB_u(t). Among them, the target state matrix is X(t), t is the current moment, deltaA is the discrete state transition matrix, the discrete input control matrix is deltaB, and the input matrix is u(t).
[0161] Then multiply the target state matrix by the output control matrix to obtain the output matrix corresponding to the current moment. That is, perform a dot product operation on the target state matrix X(t) (or expressed as X(b,d_in,n)) and the output control matrix C(b,l,n) to obtain the output matrix corresponding to the current moment, denoted as y(b,n).
[0162] Finally, determine the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment. Specifically, determine the output matrices corresponding to the input matrices at other moments in the data to be processed, stack the output matrix corresponding to the current moment and the output matrices corresponding to other moments to obtain the first target output matrix corresponding to the data to be processed. The first target output matrix is the data processing result corresponding to the data to be processed. The first target output matrix is denoted as y(b,l,n).
[0163] As can be seen from the above, in this application, by discretizing the data to be processed, it avoids performing arithmetic operations on the entire data to be processed like the first operator, which involves many processing steps in the middle. This can greatly reduce the amount of calculation. Therefore, compared with the first operator, the second operator has fewer operator units, thus having less calculation amount and faster calculation speed. And the second operator can be loaded on the central processing unit, which can reduce the calculation amount of the graphics processing unit, thereby improving the data processing speed and processing efficiency of the pre-trained model for the data to be processed as a whole.
[0164] In some embodiments, according to the second operator, data processing is performed on the data to be processed and the discretization parameters, and the data processing result includes:
[0165] (2.1) Perform data format conversion on the second operator and other operators corresponding to the pre-trained model to obtain the target pre-trained model in the target data format;
[0166] (2.2) Determine the third operator corresponding to the second operator in the target pre-trained model;
[0167] (2.3) Process the data to be processed and the discretization parameters according to the third operator to obtain a data processing result.
[0168] Among them, the data format of the pre-trained model can be converted, for example, to the target pre-trained model in the ONNX data format. During the conversion process, data format conversions of the second operator and other operators are involved.
[0169] The pre-trained model can be a VMamba model. The VMamba model is trained under a graphics processor. If you want to migrate it to other hardware processing units for inference (such as a CPU), it often needs to be first converted to the ONNX format as an intermediate, that is, the target pre-trained model, and then data processing is performed under other hardware processing units.
[0170] Then, determine the third operator corresponding to the second operator in the target pre-trained model. The third operator is actually also the corresponding operator in the ONNX data format, and its function is the same as that of the second operator. Finally, process the data to be processed and the discretization parameters according to the third operator to obtain a data processing result.
[0171] In some embodiments, processing the data to be processed and the discretization parameters according to the third operator to obtain a data processing result includes:
[0172] (2.3.1) Determine the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator;
[0173] (2.3.2) Perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain a first processing matrix, and perform one-dimensional convolution processing on the first processing matrix and the state transition matrix to obtain a target state matrix;
[0174] (2.3.3) Multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a first target input matrix;
[0175] (2.3.4) Perform block accumulation processing on the first target input matrix to obtain a second target input matrix, and determine the output matrix to be processed according to the second target input matrix, the target state matrix, and a preset state space matrix with all elements being zero;
[0176] (2.3.5) Perform one-dimensional convolution processing on the output matrix to be processed and the output control matrix to obtain a second target output matrix, and the second target output matrix is the data processing result corresponding to the data to be processed.
[0177] Among them, the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator can be understood as the operator units corresponding inside the third operator. The discrete step matrix is denoted as dts, the state transition matrix is denoted as As, the input control matrix is denoted as Bs, and the output control matrix is denoted as Cs. The input matrix corresponding to the data to be processed can also be determined, and the input matrix is denoted as Us.
[0178] These matrix data can be converted to the float data type to ensure the calculation accuracy.
[0179] Then, perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain the first processing matrix, and perform one-dimensional convolution processing on the first processing matrix and the state transition matrix to obtain the target state matrix. Specifically, the input matrix corresponding to the data to be processed can be block-processed to obtain multiple sub-matrices, and then the cumulative sum matrix corresponding to each sub-matrix can be determined according to the time step of each sub-matrix, and the cumulative sum matrices corresponding to each sub-matrix are stacked to obtain the first processing matrix.
[0180] For example, within the total time step length L corresponding to the input matrix, the input matrix is block-processed with cumsum_chunk_size as the step length to obtain multiple sub-matrices. cumsum_chunk_size is set to 6 to improve the calculation efficiency on the central processing unit. The specific representation of the sub-matrix corresponding to each time step is x_chunk = dts[cumsum_chunk_size*i, cumsum_chunk_size*(i + 1)]. Then, calculate the cumulative sum matrix corresponding to the sub-matrix of each time step, denoted as item_chunk. If the sub-matrix x_chunk is the first block, directly calculate the cumulative sum; otherwise, continue to accumulate on the basis of the cumulative sum corresponding to the previous time step. Stack the cumulative sum matrices item_chunk corresponding to each time step together to obtain the first processing matrix ts.
[0181] Then, perform one-dimensional convolution processing on the first processing matrix and the state transition matrix to obtain the target state matrix. For example, use the F.conv1d function in the PyTorch calculation library to perform one-dimensional convolution operations on the first processing matrix ts. Here, the first processing matrix ts is processed as (chunk_L*B, G*D, 1), and the state transition matrix As is processed as (-1, 1, 1). Set the number of groups groups of the convolution to G*D. This is actually performing one-dimensional convolution on the first processing matrix ts and the state transition matrix As to obtain the convolution result. Then, perform exponential processing on each element of the convolution result to obtain the state matrix Ats to be processed. Next, normalize the value of the last time step of the state matrix Ats to be processed to avoid division-by-zero errors. The specific operation is: scale = Ats[-1].detach() + 1e-30, and then divide the state matrix Ats to be processed by this scale value to obtain the normalized target state matrix rAts.
[0182] After that, multiply the input matrix, the discrete step matrix, and the input control matrix to obtain the first target input matrix. For example, perform element-wise multiplication on the input matrix Us and the discrete step matrix dts to obtain the intermediate discrete input matrix duts. Then, use the torch.matmul function to multiply the intermediate discrete input matrix duts with the input control matrix Bs to obtain the first target input matrix dtBus.
[0183] Then, perform block accumulation processing on the first target input matrix to obtain the second target input matrix, and determine the matrix to be processed output based on the second target input matrix, the target state matrix, and a preset state space matrix with all elements being zero.
[0184] Specifically, determine the sub-matrix corresponding to each time step in the first target input matrix, and accumulate the sub-matrices corresponding to each time step to obtain the second target input matrix. For example, extract the sub-matrix corresponding to each time step from the first target input matrix dtBus. The sub-matrix is expressed as x_chunk = dts[cumsum_chunk_size*i, cumsum_chunk_size*(i + 1)]. If the sub-emergency x_chunk is the sub-matrix corresponding to the first time step, directly calculate the accumulation; if not, continue to accumulate based on the accumulation result of the previous sub-matrix to obtain the second target input matrix dtBus_cumsum.
[0185] Then, perform element-wise multiplication on the target state matrix rAts and the second target input matrix dtBus_cumsum to obtain the second processing matrix hs_tmp. Multiply the second processing matrix hs_tmp with a preset state space matrix initialized with all zeros (such as the state space variable matrix hprefix matrix) to obtain the third processing matrix. Multiply the state matrix to be processed Ats with the preset state space matrix to obtain the fourth processing matrix. Finally, add the third processing matrix and the fourth processing matrix to obtain the output matrix to be processed hs.
[0186] Finally, perform one-dimensional convolution processing on the output matrix to be processed and the output control matrix to obtain the second target output matrix. The second target output matrix is the data processing result corresponding to the data to be processed. For example, use the F.conv1d function to perform one-dimensional convolution operation on the output control matrix Cs and the output matrix to be processed hs. The output control matrix Cs is processed into (1, chunk_L*B*G*N, 1), and the output matrix to be processed hs is processed into (chunk_L*B*G*D, N, 1). Set the number of groups of convolution groups to chunk_L*B*G to obtain the first output matrix ys.
[0187] If the third operator has a corresponding instruction matrix Ds, multiply the instruction matrix Ds and the input matrix us and then add the result to the first output matrix ys to obtain the final second target output matrix oys. If the third operator does not have an instruction matrix Ds, the first output matrix ys is used as the second target output matrix oys.
[0188] As can be seen from the above, by performing block processing on the input matrix and then stacking it into an overall matrix, and then performing calculations on the overall matrix, this can reduce a large number of calculation steps. Therefore, compared with the first operator, the third operator has fewer calculation units, but it can achieve the same function as the first operator, thus improving the processing speed of the data to be processed, and then improving the data processing speed and processing efficiency of the pre-trained model.
[0189] In some embodiments, in the target pre-trained model, three optimization methods, namely constant folding, dead code elimination, and operator fusion, can also be adopted to effectively reduce the computational complexity and runtime resource overhead when the target pre-trained model processes data, thereby improving the overall execution efficiency.
[0190] For example, in the constant folding optimization method, during the export process of the target pre-trained model, constant expressions inside the model such as torch.tensor([2, 3]) * torch.tensor([4, 5]) are calculated in advance as [8, 15] to avoid repeated calculations during runtime. Additionally, in the FLOPs calculation function, for expressions of static parameters (such as B, L, D, N) like flops = 9 * B * L * D * N, if the parameters are known (such as B = 1, L = 256, etc.), they can be directly folded into constant values, thereby reducing the computational amount in function calls. Moreover, for overlapping matrix operations (such as bias addition and activation functions), they are folded in advance through optimization logic before inference to further improve performance.
[0191] Another example is that through dead code elimination technology, the present application eliminates redundant logic in the code corresponding to the target pre-trained model to reduce resource waste. For example, in the main function, there are multiple branch logics for handling different tasks, such as the branch handling of args.inference_time_pytorch for calculating the inference time of the pre-trained model and args.inference_time_onnx for calculating the inference time of the target pre-trained model; if some branches are not enabled or will not be triggered (such as args.inference_time_simply_onnx for calculating the inference time of a certain special operator of the target pre-trained model is not set), the relevant logic can be directly removed through static analysis. There are also some imported modules such as selective_scan_cuda_oflex or selective_scan_cuda_core that may not be actually called, and the import logic of these unused modules can be deleted through analysis during compilation. Additionally, in some breakpoint check logics (such as pdb.set_trace()), they can be removed through dead code elimination in the production environment to avoid affecting the inference performance. By detecting and removing unused parameters, branches, and debugging code, this optimization method significantly reduces the ineffective overhead during runtime.
[0192] For another example, through the operator fusion technology, the present application combines multiple consecutive operations in the target pre-trained model into a single efficient operator for execution, further improving the inference efficiency. For example, for some consecutive tensor operations such as flatten, transpose, and flip, the original code implements these operations in multiple steps, while after optimization, they can be combined into a custom operator, custom_operator, to reduce the operation overhead on the CPU. Similarly, matrix operations such as ys[:,0:2]+ys[:,2:4].flip(dims=[-1]) and transpose can be fused through kernel fusion to reduce the storage and operation times of intermediate tensors. During the inference process of the target pre-trained model, the post-processing logic (such as normalization or activation function) after the inference output can be directly embedded into the target pre-trained model, and the integration of inference and post-processing can be achieved through the operator fusion of the target pre-trained model. In addition, in some matrix calculations, multiple torch.einsum operations, specifically torch.einsum('bdl,bdnl,bdl->bdln') and torch.einsum('bdl,dn->bdln'), can be fused by optimizing the computational graph. This optimization method of operator fusion significantly reduces the number of operator calls during runtime and the memory occupancy of intermediate tensors, which has great value for the inference deployment of the target pre-trained model.
[0193] Please refer to Figure 5 , Figure 5 which is a comparison chart of the number of operators corresponding to the pre-trained model and the target pre-trained model provided by the embodiments of the present application. Through Figure 5 it can be seen that in the present application, after optimizing the first operator and other operators in the pre-trained model, the number of operators in the target pre-trained model is significantly less, and the computational inference speed of the target pre-trained model is higher.
[0194] In the embodiments of the present application, by obtaining the pre-trained model and the model data corresponding to the pre-trained model, the pre-trained model is a model under the state space model structure; determining the first operator in the pre-trained model that needs to accelerate data processing according to the model data; decoupling the first operator to obtain the second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; obtaining the data to be processed corresponding to the second operator, and determining the discretization parameters corresponding to the data to be processed; and performing data processing on the data to be processed and the discretization parameters according to the second operator to obtain the data processing result.
[0195] Thus, by determining a first operator that requires data processing acceleration from the model data of the pre-trained model under the state space model structure, and then decoupling the first operator, a second operator can be obtained. The second operator can be applied to other hardware processing units relative to the first operator. In this way, the computing power of other hardware processing units can be utilized to accelerate the operation of the second operator. Moreover, the data to be processed by the second operator is discretized using discretization parameters. This enables the second operator to avoid directly processing the entire data to be processed, but rather processes the discretized data, reducing the complex calculations of the second operator on the overall data to be processed, thereby improving the processing speed of the second operator for the data to be processed. Therefore, compared with the limited data acceleration effect of the model acceleration technology in the related art on the model data under the state space model structure, in this application, the first operator can be decoupled to obtain the second operator, and then the second operator can be loaded by other hardware units, and the data to be processed by the second operator is discretized, which can significantly improve the data processing speed of the second operator, thus greatly improving the data processing speed of the pre-trained model of the entire state space model. Therefore, the data processing acceleration effect on the pre-trained model of the state space model is more obvious than that in the related art.
[0196] Please refer to Figure 6 , Figure 6 which is another flowchart of the model data processing method provided by the embodiment of this application. The model data processing method may include the following steps:
[0197] Step 501, obtain a pre-trained model and the model data corresponding to the pre-trained model, where the pre-trained model is a model under the state space model structure;
[0198] Step 502, determine a first operator in the pre-trained model that needs to be accelerated in data processing according to the model data;
[0199] Step 503, determine the operation logic corresponding to each sub-operator in the first operator;
[0200] Step 504, modify each sub-operator according to the pre-designed computing library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator;
[0201] Step 505, generate a second operator corresponding to the first operator according to the target sub-operator corresponding to each sub-operator;
[0202] Step 506, perform data format conversion on the second operator and other operators corresponding to the pre-trained model to obtain a target pre-trained model in the target data format;
[0203] Step 507, determine a third operator corresponding to the second operator in the target pre-trained model;
[0204] Step 508: Process the data to be processed and the discretization parameters according to the third operator to obtain a data processing result;
[0205] Step 509: Determine the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator;
[0206] Step 510: Perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain a first processed matrix, and perform one-dimensional convolution processing on the first processed matrix and the state transition matrix to obtain a target state matrix;
[0207] Step 511: Multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a first target input matrix;
[0208] Step 512: Perform block accumulation processing on the first target input matrix to obtain a second target input matrix, and determine the output matrix to be processed according to the second target input matrix, the target state matrix, and a preset state space matrix with all elements being zero;
[0209] Step 513: Perform one-dimensional convolution processing on the output matrix to be processed and the output control matrix to obtain a second target output matrix, and the second target output matrix is the data processing result corresponding to the data to be processed.
[0210] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference may be made to the detailed description of the above model data processing method, which will not be elaborated here.
[0211] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the model data processing device provided by the embodiments of the present application. This model data processing device can execute the above model data processing method.
[0212] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0213] The model data processing device 600 includes:
[0214] A first acquisition module 610, configured to acquire a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under a state space model structure;
[0215] A determination module 620, configured to determine a first operator in a pre-trained model that needs data processing acceleration according to model data;
[0216] A decoupling module 630, configured to perform decoupling processing on the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator;
[0217] A second acquisition module 640, configured to acquire data to be processed corresponding to the second operator and determine discretization parameters corresponding to the data to be processed;
[0218] A processing module 650, configured to perform data processing on the data to be processed and the discretization parameters according to the second operator to obtain a data processing result.
[0219] In some embodiments, the decoupling module 630 is configured to:
[0220] Determine the operation logic corresponding to each sub-operator in the first operator;
[0221] Determine the type of hardware processing unit corresponding to each sub-operator according to the operation logic;
[0222] Decouple each sub-operator according to the type of hardware processing unit to obtain a second operator.
[0223] In some embodiments, the decoupling module 630 is configured to:
[0224] Determine the operation logic corresponding to each sub-operator in the first operator;
[0225] Modify each sub-operator according to a pre-designed computing library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator;
[0226] Generate a second operator corresponding to the first operator according to the target sub-operator corresponding to each sub-operator.
[0227] In some embodiments, the processing module 650 is configured to:
[0228] Determine a state transition matrix, an input control matrix, and an output control matrix corresponding to the second operator;
[0229] Determine a discrete step matrix according to the discretization parameters, multiply the discrete step matrix by the state transition matrix, and perform exponentiation processing on each element in the multiplied matrix to obtain a discrete state transition matrix;
[0230] Determine an input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a discrete input control matrix;
[0231] Obtain the state matrix corresponding to the previous moment, multiply the state matrix by the discrete state transition matrix, and then add the discrete input control matrix to obtain the target state matrix corresponding to the current moment;
[0232] Multiply the target state matrix by the output control matrix to obtain the output matrix corresponding to the current moment;
[0233] Determine the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment.
[0234] In some embodiments, the processing module 650 is configured to:
[0235] Determine the output matrix corresponding to the input matrix at other moments in the data to be processed;
[0236] Perform matrix stacking processing on the output matrix corresponding to the current moment and the output matrices corresponding to other moments to obtain the first target output matrix corresponding to the data to be processed, and the first target output matrix is the data processing result corresponding to the data to be processed.
[0237] In some embodiments, the processing module 650 is configured to:
[0238] Perform data format conversion on the second operator and other operators corresponding to the pre-trained model to obtain the target pre-trained model in the target data format;
[0239] Determine the third operator corresponding to the second operator in the target pre-trained model;
[0240] Perform data processing on the data to be processed and the discretization parameters according to the third operator to obtain the data processing result.
[0241] In some embodiments, the processing module 650 is configured to:
[0242] Determine the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator;
[0243] Perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain the first processing matrix, and perform one-dimensional convolution processing on the first processing matrix and the state transition matrix to obtain the target state matrix;
[0244] Multiply the input matrix, the discrete step matrix, and the input control matrix to obtain the first target input matrix;
[0245] Perform block accumulation processing on the first target input matrix to obtain the second target input matrix, and determine the output matrix to be processed according to the second target input matrix, the target state matrix, and the preset state space matrix with all elements being zero;
[0246] Perform a one-dimensional convolution process on the output matrix to be processed and the output control matrix to obtain a second target output matrix, which is the data processing result corresponding to the data to be processed.
[0247] In some embodiments, the processing module 650 is configured to:
[0248] Perform a block processing on the input matrix corresponding to the data to be processed to obtain a plurality of sub-matrices;
[0249] Determine the cumulative sum matrix corresponding to each sub-matrix according to the time step of each sub-matrix, and perform a matrix stacking process on the cumulative sum matrices corresponding to each sub-matrix to obtain a first processing matrix.
[0250] In some embodiments, the processing module 650 is configured to:
[0251] Determine the sub-matrix corresponding to each time step in the first target input matrix;
[0252] Accumulate the sub-matrices corresponding to each time step to obtain a second target input matrix.
[0253] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the detailed description of the above model data processing method, which will not be elaborated here.
[0254] In the embodiments of the present application, the first acquisition module 610 acquires a pre-trained model and model data corresponding to the pre-trained model, and the pre-trained model is a model under a state space model structure; the determination module 620 determines a first operator in the pre-trained model that needs to accelerate data processing according to the model data; the decoupling module 630 performs a decoupling process on the first operator to obtain a second operator corresponding to the first operator, and the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; the second acquisition module 640 acquires the data to be processed corresponding to the second operator and determines the discretization parameter corresponding to the data to be processed; the processing module 650 performs data processing on the data to be processed and the discretization parameter according to the second operator to obtain a data processing result.
[0255] Thus, a first operator that requires data processing acceleration is determined from the model data of the pre-trained model under the state space model structure, and then the first operator is decoupled to obtain a second operator. The second operator can be applied to other hardware processing units relative to the first operator. In this way, the computing power of other hardware processing units can be utilized to accelerate the operation of the second operator. Moreover, the data to be processed by the second operator is discretized using discretization parameters. This enables the second operator to avoid directly processing the data to be processed as a whole and instead process the discretized data, reducing the complex calculations of the second operator on the overall data to be processed and thus improving the processing speed of the second operator for the data to be processed. Therefore, compared with the limited data acceleration effect of the model acceleration technology on the model data under the state space model structure in the related art, in this application, the first operator can be decoupled to obtain the second operator, and then the second operator can be loaded by other hardware units, and the data to be processed by the second operator is discretized, which can significantly improve the data processing speed of the second operator and thus greatly improve the data processing speed of the pre-trained model of the entire state space model. Therefore, the data processing acceleration effect on the pre-trained model of the state space model is more obvious than that in the related art.
[0256] An embodiment of this application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned model data processing method is implemented. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0257] Please refer to Figure 8 , Figure 8 which shows the hardware structure of a computer device in another embodiment. The computer device includes:
[0258] A processor 701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0259] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702, and the processor 701 is used to call and execute the model data processing method of the embodiments of this application;
[0260] The input / output interface 703 is used to implement information input and output;
[0261] The communication interface 704 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0262] The bus 705 transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);
[0263] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.
[0264] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned model data processing method is implemented.
[0265] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0266] The model data processing method, model data processing device, computer device, and storage medium provided by the embodiments of the present application obtain a pre-trained model and model data corresponding to the pre-trained model through, in the embodiments of the present application. The pre-trained model is a model under a state space model structure; determine a first operator in the pre-trained model that needs to be accelerated in data processing according to the model data; perform decoupling processing on the first operator to obtain a second operator corresponding to the first operator, and the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; obtain the data to be processed corresponding to the second operator, and determine the discretization parameter corresponding to the data to be processed; perform data processing on the data to be processed and the discretization parameter according to the second operator to obtain a data processing result.
[0267] Therefore, by determining the first operator that needs to be accelerated in data processing from the model data of the pre-trained model under the state space model structure, and then performing decoupling processing on the first operator, the second operator can be obtained. The second operator can be applied to other hardware processing units relative to the first operator. In this way, the computing power of other hardware processing units can be utilized to accelerate the operation of the second operator. Moreover, the data to be processed that needs to be processed by the second operator is discretized using the discretization parameter. This enables the second operator to avoid directly processing the data to be processed as a whole, but rather processes the discretized data. This can reduce the complex calculations of the second operator on the overall data to be processed, thereby improving the processing speed of the second operator for the data to be processed. Therefore, compared with the limited data acceleration processing effect of the model acceleration technology in the related art for the model under the state space model structure, in the present application, the first operator can be decoupled to obtain the second operator, and then the second operator can be loaded by other hardware units, and the data to be processed of the second operator is discretized, which can significantly improve the data processing speed of the second operator, thereby significantly improving the data processing speed of the pre-trained model of the entire state space model. Therefore, the data processing acceleration effect on the pre-trained model of the state space model is more obvious than that in the related art.
[0268] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0269] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0270] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0271] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0272] As used in the specification of this application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0273] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0274] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0275] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0276] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0277] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM for short), random access memory (RAM for short), magnetic disks, or optical discs that can store programs.
[0278] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.
Claims
1. A method for processing model data, characterized in that, Including: Obtain a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under a state space model structure; Determine a first operator in the pre-trained model that needs to be accelerated in data processing according to the model data; Perform decoupling processing on the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; Obtain the data to be processed corresponding to the second operator and determine the discretization parameters corresponding to the data to be processed; Determine the state transition matrix, input control matrix, and output control matrix corresponding to the second operator; Determine a discrete step matrix according to the discretization parameters, multiply the discrete step matrix by the state transition matrix, and perform exponentiation processing on each element in the multiplied matrix to obtain a discrete state transition matrix; Determine an input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a discrete input control matrix; Obtain a state matrix corresponding to the previous moment, multiply the state matrix by the discrete state transition matrix and add the discrete input control matrix to obtain a target state matrix corresponding to the current moment; Multiply the target state matrix by the output control matrix to obtain an output matrix corresponding to the current moment; Determine the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment.
2. The model data processing method according to claim 1, wherein The performing decoupling processing on the first operator to obtain a second operator corresponding to the first operator includes: Determine the operation logic corresponding to each sub-operator in the first operator; Determine the type of hardware processing unit corresponding to each sub-operator according to the operation logic; Decouple each sub-operator according to the type of hardware processing unit to obtain a second operator.
3. The model data processing method according to claim 1, wherein The performing decoupling processing on the first operator to obtain a second operator corresponding to the first operator includes: Determine the operation logic corresponding to each sub-operator in the first operator; Modify each sub-operator according to a pre-designed library and the operation logic corresponding to each sub-operator to generate a target sub-operator corresponding to each sub-operator; Generate a second operator corresponding to the first operator according to the target sub-operator corresponding to each sub-operator.
4. The model data processing method according to claim 1, wherein The determining the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment includes: Determine the output matrix corresponding to the input matrix at other moments in the data to be processed; Perform matrix stacking processing on the output matrix corresponding to the current moment and the output matrix corresponding to other moments to obtain a first target output matrix corresponding to the data to be processed, and the first target output matrix is the data processing result corresponding to the data to be processed.
5. The model data processing method according to claim 3, wherein After the obtaining the data to be processed corresponding to the second operator and determining the discretization parameters corresponding to the data to be processed, it further includes: Perform data format conversion on the second operator and other operators corresponding to the pre-trained model to obtain a target pre-trained model in a target data format; Determine the third operator corresponding to the second operator in the target pre-trained model; Perform data processing on the data to be processed and the discretization parameter according to the third operator to obtain a data processing result.
6. The model data processing method according to claim 5, wherein The performing data processing on the data to be processed and the discretization parameter according to the third operator to obtain a data processing result includes: Determine the discrete step matrix, state transition matrix, input control matrix, and output control matrix corresponding to the third operator; Perform block cumulative processing on the input matrix corresponding to the data to be processed to obtain a first processed matrix, and perform one-dimensional convolution processing on the first processed matrix and the state transition matrix to obtain a target state matrix; Multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a first target input matrix; Perform block accumulation processing on the first target input matrix to obtain a second target input matrix, and determine a matrix to be processed for output according to the second target input matrix, the target state matrix, and a preset state space matrix with all elements being zero; Perform one-dimensional convolution processing on the matrix to be processed for output and the output control matrix to obtain a second target output matrix, and the second target output matrix is the data processing result corresponding to the data to be processed.
7. The model data processing method according to claim 6, wherein The performing block cumulative processing on the input matrix corresponding to the data to be processed to obtain a first processed matrix includes: Perform block processing on the input matrix corresponding to the data to be processed to obtain a plurality of sub-matrices; Determine the cumulative sum matrix corresponding to each sub-matrix according to the time step of each sub-matrix, and perform matrix stacking processing on the cumulative sum matrices corresponding to each sub-matrix to obtain a first processed matrix.
8. The model data processing method according to claim 6, wherein The performing block accumulation processing on the first target input matrix to obtain a second target input matrix includes: Determine the sub-matrix corresponding to each time step in the first target input matrix; Accumulate the sub-matrices corresponding to each time step to obtain a second target input matrix.
9. A model data processing device, characterized in that, including: A first acquisition module, configured to acquire a pre-trained model and model data corresponding to the pre-trained model, where the pre-trained model is a model under a state space model structure; A determination module, configured to determine a first operator that needs to perform data processing acceleration in the pre-trained model according to the model data; A decoupling module, configured to perform decoupling processing on the first operator to obtain a second operator corresponding to the first operator, where the hardware processing unit corresponding to the second operator is different from the hardware processing unit corresponding to the first operator; A second acquisition module, configured to acquire data to be processed corresponding to the second operator and determine a discretization parameter corresponding to the data to be processed; A processing module, configured to determine a state transition matrix, an input control matrix, and an output control matrix corresponding to the second operator; Determine a discrete step matrix according to the discretization parameter, multiply the discrete step matrix and the state transition matrix, and perform exponential processing on each element in the multiplied matrix to obtain a discrete state transition matrix; Determine the input matrix corresponding to the current moment according to the data to be processed, and multiply the input matrix, the discrete step matrix, and the input control matrix to obtain a discrete input control matrix; Obtain the state matrix corresponding to the previous moment, multiply the state matrix by the discrete state transition matrix, and then add the result to the discrete input control matrix to obtain the target state matrix corresponding to the current moment; Multiply the target state matrix by the output control matrix to obtain the output matrix corresponding to the current moment; Determine the data processing result corresponding to the data to be processed according to the output matrix corresponding to the current moment.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the model data processing method according to any one of claims 1 to 8.
11. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the model data processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Decoupling state spatial prediction control method of chemical multivariate processes
CN102902201A
Model conversion method and device
CN112819153A