Data processing method and device, electronic equipment, storage medium and program product
By dividing the input data into multiple sub-segments and utilizing the contextual relationship of the static computation graph, combined with hardware resources and inference time, the segmentation combination of the static computation graph is optimized, which solves the low processing efficiency problem of the fixed input length of the static computation graph in the terminal device, and achieves faster inference speed and higher data processing accuracy.
Patent Information
- Application Number
- CN202510661093.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-10-10
AI Technical Summary
In the existing technology, the input length of the static computation graph of the terminal device is fixed, which makes it difficult to meet the needs of various practical applications and has low processing efficiency. In particular, it is impossible to effectively reason about extremely long input data.
The input data is divided into multiple sub-segments, and each sub-segment is used as input to execute the static computation graph corresponding to each sub-segment. The execution result of the previous static computation graph is used as the context of the next static computation graph. The optimal segmentation combination is determined by combining hardware resources and the inference time of the static computation graph to adapt to different data scenarios.
It improves the reasoning speed and efficiency of static computation graphs, solves the problem of being unable to reason with extremely long input data, and ensures the accuracy and efficiency of data processing.
Smart Images

Figure CN120764665A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art
[0002] With the rapid development of large model technology, more and more large models are being deployed on edge devices, such as mobile phones and vehicles. The NPU (Neural Network Processing Unit) of edge devices generally uses static computational graph reasoning to achieve higher inference efficiency.
[0003] The input length of static computation graphs is fixed during the model building phase, which makes it difficult to meet the needs of various practical applications. Summary of the Invention
[0004] The present disclosure provides a data processing method, device, electronic device, storage medium and program product to solve the problem of low processing efficiency in related technologies.
[0005] According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, which includes: receiving input data; dividing the input data into multiple sub-segments; using the multiple sub-segments as input to execute a static computation graph corresponding to each sub-segment; and outputting an execution result of the static computation graph.
[0006] By dividing the input data into multiple sub-segments, the sub-segments can be executed using static computation graphs with shorter input lengths, resulting in faster and more efficient inference. Furthermore, dividing the input data into multiple sub-segments solves the problem of being unable to infer overly long input data.
[0007] In some exemplary embodiments of the present disclosure, taking the multiple sub-segments as input respectively to execute the static computation graph corresponding to each sub-segment includes: taking the multiple sub-segments as input in sequence, and executing the static computation graph corresponding to each sub-segment in sequence, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
[0008] Using the execution result of the previous static computation graph as the context of the next static computation graph can better capture and understand the contextual information of the entire input data, rather than processing each sub-segment in isolation, thereby improving the accuracy of the execution results.
[0009] In some exemplary embodiments of the present disclosure, dividing the input data into a plurality of sub-segments includes: dividing the input data into a plurality of sub-segments according to the length of the input data.
[0010] Splitting the input data according to its length ensures that each sub-segment matches the appropriate input length of the static computation graph, allowing the input data to be processed efficiently without losing information.
[0011] In some exemplary embodiments of the present disclosure, dividing the input data into a plurality of sub-segments according to the length of the input data includes: dividing the input data into a plurality of sub-segments according to the same step size.
[0012] Since the step size is the same, the starting and ending points of all sub-segments follow a unified standard, avoiding data deviation caused by random or irregular segmentation.
[0013] In some exemplary embodiments of the present disclosure, dividing the input data into multiple sub-segments according to the length of the input data includes: dividing the input data into multiple sub-segments according to the length of the input data, the multiple sub-segments including sub-segments of different lengths.
[0014] After the input data is divided into multiple sub-segments, the lengths of the sub-segments can be different to accommodate more data processing scenarios.
[0015] In some exemplary embodiments of the present disclosure, dividing the input data into multiple sub-segments according to the length of the input data includes: dividing the input data into multiple sub-segments based on the length of the input data and the input length of each static computation graph.
[0016] By segmenting the input data according to the input length of the static computation graph, the length of the segmented sub-segments can be made as consistent as possible with the input length of the static computation graph, minimizing the padding of the sub-segments and further reducing the calculation of the padding part by the static computation graph, thereby further improving the inference speed.
[0017] In some exemplary embodiments of the present disclosure, dividing the input data into multiple sub-segments based on the length of the input data and the input length of each static computation graph includes: determining a target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph; and dividing the input data into multiple sub-segments based on the input length of each target static computation graph in the target static computation graph combination.
[0018] Based on the inference time of the static computation graph, the optimal segmentation combination can be found to further improve the inference speed of the input data.
[0019] In some exemplary embodiments of the present disclosure, determining the target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph includes: determining the target static computation graph combination based on the length of the input data, the input length of each static computation graph, the inference time of each static computation graph, and the switching time of each static computation graph.
[0020] In the process of determining the optimal segmentation combination, the impact of the switching time of the static computation graph on the total inference time is taken into account to further improve the inference speed of the input data.
[0021] In some exemplary embodiments of the present disclosure, the inference time of each static computation graph is determined by the hardware resources on which the static computation graph is deployed.
[0022] When deploying static computation graphs, consider the impact of underlying hardware capabilities on the execution efficiency of the static computation graphs, deploy static computation graphs with appropriate input lengths, and further improve data inference efficiency.
[0023] In some exemplary embodiments of the present disclosure, the various static computation graphs share model weights.
[0024] Multiple static computation graphs share the same model parameters, avoiding the overhead of storing parameters separately for each static computation graph, thereby saving memory.
[0025] In some exemplary embodiments of the present disclosure, the method further includes:
[0026] The last sub-segment of the multiple sub-segments is padded to match the input length of the corresponding static computation graph.
[0027] By padding the sub-segments, reasoning can be performed on sub-segments whose length is less than the input length of the static computation graph to adapt to different sub-segment lengths.
[0028] According to a second aspect of an embodiment of the present disclosure, a data processing device is provided, which includes: a receiving module for receiving input data; a segmentation module for dividing the input data into multiple sub-segments; an execution module for using the multiple sub-segments as input to execute a static computation graph corresponding to each sub-segment; and an output module for outputting the execution result of the static computation graph.
[0029] In some exemplary embodiments of the present disclosure, the execution module is specifically configured to take the multiple sub-segments as input in sequence and execute the static computation graphs corresponding to the sub-segments in sequence, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
[0030] In some example embodiments of the present disclosure, the splitting module is specifically configured to split the input data into a plurality of sub-segments according to a length of the input data.
[0031] In some example embodiments of the present disclosure, the splitting module is specifically configured to split the input data into a plurality of sub-segments according to a length of the input data.
[0032] In some example embodiments of the present disclosure, the splitting module is specifically configured to split the input data into a plurality of sub-segments according to a length of the input data, wherein the plurality of sub-segments include sub-segments with different lengths.
[0033] In some example embodiments of the present disclosure, the splitting module is specifically configured to split the input data into a plurality of sub-segments based on a length of the input data and input lengths of the respective static computation graphs.
[0034] In some example embodiments of the present disclosure, the splitting module includes: a combination determining unit configured to determine a target static computation graph combination based on a length of the input data, input lengths of the respective static computation graphs, and inference times of the respective static computation graphs; and a splitting unit configured to split the input data into a plurality of sub-segments based on the input lengths of the respective static computation graphs in the target static computation graph combination.
[0035] In some example embodiments of the present disclosure, the combination determining unit is specifically configured to determine a target static computation graph combination based on a length of the input data, input lengths of the respective static computation graphs, inference times of the respective static computation graphs, and switching times of the respective static computation graphs.
[0036] In some example embodiments of the present disclosure, the inference times of the respective static computation graphs are determined by hardware resources on which the static computation graphs are deployed.
[0037] In some example embodiments of the present disclosure, the respective static computation graphs share model weights.
[0038] In some example embodiments of the present disclosure, the apparatus further includes a padding module configured to pad a last sub-segment of the plurality of sub-segments to match an input length of a corresponding static computation graph.
[0039] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects:
[0040] Splitting the input data into a plurality of sub-segments and executing the sub-segments by corresponding static computation graphs can achieve a faster inference speed and higher efficiency than directly executing a static computation graph with a longer input length on all input data.
[0041] Further, splitting the input data can solve the problem of inability to infer on super-long input data.
[0042] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0044] Figure 1 is a flowchart illustrating a data processing method according to some embodiments of the present disclosure.
[0045] Figure 2 is a flowchart illustrating a data processing method according to other embodiments of the present disclosure.
[0046] Figure 3 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0047] Figure 4 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0048] Figure 5 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0049] Figure 6 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0050] Figure 7 This is a schematic diagram showing different inference times corresponding to different input lengths according to some embodiments of the present disclosure.
[0051] Figure 8 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0052] Figure 9 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0053] Figure 10 is a flowchart of a data processing method according to some further embodiments of the present disclosure.
[0054] Figure 11 It is a flowchart of an implementation of a data processing method according to some embodiments of the present disclosure.
[0055] Figure 12 This is a scenario diagram of the application of the data processing method in a terminal according to some embodiments of the present disclosure.
[0056] Figure 13 It is a scene diagram of the application of the data processing method in a vehicle according to some other embodiments of the present disclosure.
[0057] Figure 14 It is a block diagram of a data processing device according to some embodiments of the present disclosure.
[0058] Figure 15 is a block diagram of a data processing device according to some further embodiments of the present disclosure.
[0059] Figure 16 is a block diagram of a data processing device according to some further embodiments of the present disclosure.
[0060] Figure 17 The figure is a schematic diagram of an electronic device shown in an exemplary embodiment. DETAILED DESCRIPTION
[0061] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.
[0062] First, the terms involved in the embodiments of the present disclosure are explained.
[0063] A static computation graph is a computational process that is predefined before execution. It usually consists of a series of fixed computational operations (such as matrix multiplication, activation functions, etc.), and the order and structure of the above computational operations are determined before the static computation graph is run.
[0064] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0065] The data processing method is explained below with specific embodiments and accompanying drawings.
[0066] Figure 1This is a flowchart illustrating a data processing method according to some embodiments of the present disclosure. This data processing method can be deployed and executed on a cloud server and / or on a terminal device, which can include various terminal devices such as vehicles and mobile phones. For example, in the field of smart cockpits, using the vehicle as a terminal device to perform various tasks, such as intelligent voice assistants and personalized services, can provide faster response times and a better user experience.
[0067] like Figure 1 As shown, the data processing method includes the following steps.
[0068] In step S102 , input data is received.
[0069] Input data refers to data that needs to be input into the static computation graph for inference. Input data may include, but is not limited to, text data, image data, audio data, video data, and so on.
[0070] In some embodiments, for text data, the text data is segmented to obtain multiple tokens (words), and the multiple tokens are embedded to obtain an embedding vector for the text data. For image data, a pre-trained image model is used to extract image features of the image data, and the image features are embedded to obtain an embedding vector for the image data. For audio data, the audio signal is converted into a spectrogram, and a pre-trained audio model is used to extract features from the spectrogram to obtain audio features, and the audio features are embedded to obtain an embedding vector for the audio data. For video data, a pre-trained video model is used to extract features from the video data to obtain video features, and the video features are embedded to obtain an embedding vector for the video data.
[0071] In some exemplary embodiments of the present disclosure, the input data is text data as an example, and the input data has a certain length. The length of the text data is expressed by the number of tokens, the length of the video data is expressed by the height, width and number of color channels of the image, the length of the audio data is expressed by the time length of the audio data or the number of audio frames included in the audio data, and the length of the video data is expressed by the number of frames included in the video clip, and the height, width and number of color channels of each video frame.
[0072] In some exemplary embodiments of the present disclosure, in a single-round question-and-answer scenario, the input data may include the prompt word entered by the user in the current round of the intelligent question-and-answer system. In a multi-round text scenario, the input data may include the prompt word entered by the user in the current round of the intelligent question-and-answer system, the prompt words entered in all previous rounds of the current round of the question-and-answer system, and the result data generated in all previous rounds of the current round of the question-and-answer system.
[0073] In some embodiments, in a language model based on a Transformer architecture, corresponding Key and Value vectors are generated for input data during a self-attention mechanism calculation process of the language model on the input data. When a Key Value (KV) Cache technology is used, the Key and Value are not discarded but stored in a cache. When the next input data is input to the language model, the language model uses the Key and Value in the previous cache for attention calculation in addition to calculating new Key and Value for the new input data, which can effectively reduce repeated calculation and improve inference efficiency. Therefore, in the present embodiment, the input data includes not only the input data received in the present round but also the cached Key-Value matrix corresponding to the input data processed before the present round.
[0074] Receiving input data includes: receiving input data input by a user in the present round, obtaining a Key-Value matrix corresponding to input data processed before the present round from a cache, performing tokenization on the input data in the present round to obtain a Token sequence, merging the Token sequence with the Key-Value matrix corresponding to the input data processed before the present round, and obtaining input data finally input to a static calculation graph for processing.
[0075] In step S104, the input data is divided into multiple sub-segments.
[0076] Each sub-segment refers to a continuous part divided from the input data.
[0077] In some embodiments, dividing the input data into multiple sub-segments can include: dividing the input data according to a fixed length to obtain multiple sub-segments with the same length. The fixed length can be set according to actual conditions. It should be noted that if the input data is divided according to a fixed length, the length of the last sub-segment may be less than the fixed length, in which case the last sub-segment can be padded to obtain a sub-segment with a fixed length.
[0078] In one possible implementation, the input data is divided into multiple sub-segments according to the length of the input data.
[0079] The length of the input data can be understood as the data size of the input data. In this embodiment, the length of the input data is represented by the number of tokens included in the input data. In other words, the length of the input data is the number of tokens included in the input data. For example: if the input data includes 1000 tokens, then the length of the input data is 1000. It should be noted that the representation method of the length of the input data is not limited. In practical applications, it can also be represented by the number of characters, tensor size, etc.
[0080] In some exemplary embodiments of the present disclosure, the input data is text data as an example for description. The text is segmented, that is, the input data is decomposed into tokens to obtain a token list, which is then divided into multiple sub-segments.
[0081] For example, if each sub-segment contains at most 6 tokens, 18 tokens can be divided into multiple sub-segments, with each sub-segment being no longer than 6 tokens. Another example is to split the token list into two quantiles, dividing the 18 tokens into two sub-segments, with each sub-segment being 9 tokens long.
[0082] For example, 18 tokens are divided into two sub-segments, one with a length of 10 tokens and the other with a length of 8 tokens. The length of the last sub-segment (8 tokens) is less than the fixed length of 10 tokens. In this case, the last sub-segment can be padded with 2 tokens to obtain a fixed-length sub-segment.
[0083] In this embodiment, by segmenting the input data according to its length, it is possible to ensure that each sub-segment matches the appropriate input length of the static computation graph, so that the input data can be processed efficiently without losing information.
[0084] In step S106 , the static computation graph corresponding to each sub-segment is executed by taking the multiple sub-segments as input.
[0085] In the embodiment of the present application, each sub-segment has a corresponding static computation graph, and after the sub-segment is input into the corresponding static computation graph, the static computation graph is executed.
[0086] Each sub-segment is passed as input to its corresponding static computation graph, and the computation logic predefined in the static computation graph is executed, and the static computation graph generates the execution result.
[0087] In step S108, the execution result of the static computation graph is output.
[0088] In some exemplary embodiments of the present disclosure, the execution results of multiple static computation graphs are integrated or fused to generate and output the execution results of the input data.
[0089] In some exemplary embodiments of the present disclosure, the execution results of multiple static computation graphs are directly concatenated to form a larger feature vector, which is used as the inference result of the input data. For example, if two static computation graphs output vectors of length d1 and d2, respectively, they can be concatenated into a vector of length d1 + d2, which is then output as the inference result of the input data.
[0090] In some exemplary embodiments of the present disclosure, a weighted sum is performed on the execution results of multiple static computation graphs to generate an execution result for the input data. The weights can be fixed or obtained through training. Weights can also be dynamically assigned to the output of each static computation graph using an attention mechanism.
[0091] In some exemplary embodiments of the present disclosure, the execution result of the previous static computation graph is passed to the next static computation graph according to the order of the sub-segments in the input data, and the execution result of the last static computation graph is used as the inference result of the input data.
[0092] In this embodiment, after receiving input data, the input data is divided into multiple sub-segments. Based on the multiple sub-segments and the static computation graphs corresponding to each sub-segment, the multiple sub-segments are used as input to execute the corresponding static computation graphs, and finally the execution results of the static computation graphs are output. Dividing the input data into multiple sub-segments allows the sub-segments to be executed with static computation graphs with fixed input lengths. Compared to the static computation graphs with longer input lengths in the prior art, the inference speed is faster and more efficient. In addition, the input data can be split, solving the problem of inference failure with extremely long input data.
[0093] like Figure 2 As shown, it is a flowchart of the implementation process of the data processing method provided by some other exemplary embodiments of the present disclosure, which includes the following steps.
[0094] S202: Receive input data.
[0095] S204: Divide the input data into multiple sub-segments.
[0096] S206: Take the multiple sub-segments as input in sequence and execute the static computation graph corresponding to each sub-segment in order, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
[0097] The above order can be understood as the order of the sub-segments in the input data. For example, if the input data includes sub-segment 1, sub-segment 2, and sub-segment 3, and there is a sequence from sub-segment 1 to sub-segment 2, and then to sub-segment 3, then the static computation graph corresponding to sub-segment 1 will be executed first using sub-segment 1 as input, then the static computation graph corresponding to sub-segment 2 will be executed using sub-segment 2 as input, and finally the static computation graph corresponding to sub-segment 3 will be executed using sub-segment 3 as input.
[0098] In computer science, context can be understood as the state or environment associated with a program during execution. In this embodiment, the execution result of the previous static computation graph is used as context and stored in the KV cache. When the next sub-segment is input into the subsequent static computation graph and executed, the execution result of the previous static computation graph is read from the KV cache as context to more accurately understand the meaning of the next sub-segment.
[0099] In some exemplary embodiments of the present disclosure, the input data includes sub-segment 1, sub-segment 2, and sub-segment 3. There is a sequence from sub-segment 1 to sub-segment 2 and then to sub-segment 3 in the input data. Accordingly, sub-segment 1 is first used as input to execute the static computation graph A corresponding to sub-segment 1, and the computation result of the static computation graph A is obtained and stored in the KV cache. Then, sub-segment 2 is input into the static computation graph B, and the computation result of the static computation graph A is read from the KV cache. The static computation graph B corresponding to sub-segment 2 is executed, and the computation result of the static computation graph B is obtained and stored in the KV cache. Then, sub-segment 3 is input into the static computation graph C, and the computation result of the static computation graph B is read from the KV cache. The static computation graph B corresponding to sub-segment 2 is executed, and the computation result of the static computation graph C is obtained.
[0100] S208: Output the execution result of the static computation graph.
[0101] The execution result of the static computation graph corresponding to the last sub-segment in S206 is output as the inference result of the input data.
[0102] In this embodiment, using the execution result of the previous static computation graph as the context of the next static computation graph can better capture and understand the contextual information of the entire input data, rather than processing each sub-segment in isolation, thereby improving analysis accuracy.
[0103] like Figure 3 As shown in FIG, it is a flowchart of the implementation process of the data processing method provided by some further exemplary embodiments of the present disclosure, which includes the following steps.
[0104] In step S302 , input data is received.
[0105] In step S304 , the input data is divided into multiple sub-segments according to the same step size.
[0106] The same step size can be understood as that the distance from the starting point of the previous sub-segment to the starting point of the next sub-segment remains consistent each time a new sub-segment is intercepted.
[0107] Divide the input data into multiple sub-segments using the same step size, including: Dividing the input data into multiple sub-segments using the same step size based on the length of the input data. For example, if the input data is a sequence of tokens and the step size is 5 tokens, then when moving to the next sub-segment, the starting point of the new sub-segment is determined by moving 5 tokens backward from the start of the previous sub-segment. Following this process, the input data is divided until the number of sub-segments is less than or equal to the step size.
[0108] For example, if the input data is 18 tokens long and has a step length of 5 tokens, the segmentation results are: sub-segment 1 includes 5 tokens, sub-segment 2 includes 5 tokens, sub-segment 3 includes 5 tokens, and sub-segment 4 includes 3 tokens. The length of sub-segment 4 is 3 tokens less than the step length of 5 tokens. In this case, sub-segment 4 can be padded with 2 tokens to obtain sub-segments of the same step length.
[0109] In step S306 , the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment.
[0110] In step S308, the execution result of the static computation graph is output.
[0111] In this embodiment, since the step lengths are the same, the starting points and ending points of all sub-segments follow a unified standard, thus avoiding data deviation caused by random or irregular segmentation.
[0112] like Figure 4 As shown in FIG, it is a flowchart of the implementation process of the data processing method provided by some further exemplary embodiments of the present disclosure, which includes the following steps.
[0113] In step S402 , input data is received.
[0114] In step S404, the input data is divided into a plurality of sub-segments according to the length of the input data, and the plurality of sub-segments include sub-segments of different lengths.
[0115] The length of a sub-segment may be understood as the length of the data included in the sub-segment. For example, the length of a sub-segment may be the number of tokens included in the sub-segment.
[0116] The multiple subsegments including subsegments of different lengths can be understood as the multiple subsegments not having the same length. In other words, the length of at least one subsegment among the multiple subsegments is different from the lengths of the other subsegments. Exemplarily, the subsegments of different lengths include multiple subsegments having different lengths, for example, three subsegments having lengths of 3, 4, and 5, respectively. Alternatively, the subsegments of different lengths include at least one subsegment having a different length, for example, three subsegments having lengths of 5, 4, and 5, respectively.
[0117] In some exemplary embodiments of the present disclosure, the input data is divided into several sub-segments based on the length of the input data, and the lengths of these sub-segments can be different. For example, the input data is a string of "123456789". According to a certain rule, it is divided into sub-segments of different lengths. The segmentation results are as follows: Sub-segment 1: "12", length 2; Sub-segment 2: "345", length 3; Sub-segment 3: "6789", length 4. Or another segmentation method, the segmentation results are as follows: Sub-segment 1: "1", length 1; Sub-segment 2: "2345", length 4; Sub-segment 3: "6789", length 4.
[0118] In some exemplary embodiments of the present disclosure, if the input data is divided into multiple sub-segments according to the same step size, then when the length of the input data cannot be divided by the step size, the length of the last sub-segment may be different from the lengths of other sub-segments.
[0119] In step S406, the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment;
[0120] In step S408, the execution result of the static computation graph is output.
[0121] In this embodiment, after the input data is divided into multiple sub-segments, the lengths of the multiple sub-segments may be different to accommodate more processing scenarios.
[0122] like Figure 5 As shown in FIG, it is a flowchart of the implementation process of the data processing method provided by some further exemplary embodiments of the present disclosure, which includes the following steps.
[0123] In step S502 , input data is received.
[0124] In step S504 , the input data is divided into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph.
[0125] In some exemplary embodiments of the present disclosure, in step S404, the step length may be determined based on the input length of the static computation graph deployed in the processor. The step length is any one of the input lengths of the static computation graph deployed in the processor. For example, if three static computation graphs with input lengths of 32, 128, and 384 are deployed in the processor, the step length may be any one of 32, 128, and 384.
[0126] In some exemplary embodiments of the present disclosure, the length of the input data is compared with the input length of each static computation graph. If the length of the input data is greater than the maximum input length, the maximum input length is used as the step length to divide the input data into multiple sub-segments.
[0127] For example, the maximum input length of the static computation graph is 384, and the input data is [t1, t2, t3, ..., t1000]. The segmentation is performed according to the maximum input length of 384, and the segmentation results are: sub-segment 1: [t1, t2, ..., t384]; sub-segment 2: [t385, t386, ..., t768]; sub-segment 3: [t769, t770, ..., t1000].
[0128] It should be noted that the maximum input length is used as an example for explanation, and the input length of other static computation graphs can also be used to divide the input data into multiple sub-segments.
[0129] In some exemplary embodiments of the present disclosure, the input data is divided into multiple sub-segments according to the input length of each static computation graph. In other words, the length of the sub-segments after segmentation is made to completely match the input length of the static computation graph. For example: the length of the input data is 512 tokens, and the input length of the static computation graph is 384 and 128. The segmentation is performed in a manner such that the sub-segments after segmentation completely match the input length of the static computation graph. The segmentation results are: Sub-segment 1: 384 tokens, Sub-segment 2: 128 tokens. Alternatively, Sub-segment 1: 128 tokens, Sub-segment 2: 384 tokens.
[0130] In step S506 , the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment.
[0131] In step S508, the execution result of the static calculation graph is output
[0132] In this embodiment, the input data is segmented according to the input length of the static computation graph, so that the length of the segmented sub-segments can be as close as possible to the input length of the static computation graph, thereby minimizing the padding of the sub-segments and further reducing the calculation of the padding part by the static computation graph, thereby further improving the inference speed.
[0133] like Figure 6 As shown, it is a flowchart of the implementation process of step S504 provided in some exemplary embodiments of the present disclosure, including the following steps.
[0134] In step S602 , input data is received.
[0135] In step S604, the input data is divided into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph, wherein the inference time of each static computation graph is determined by the hardware resources on which the static computation graph is deployed.
[0136] The input length of each static computation graph refers to the input length of the static computation graph deployed in the processor. The inference time of the static computation graph can be understood as the time required to execute the static computation graph. The inference time of the static computation graph can be expressed as Figure 7 The line chart shown is OK.
[0137] In some embodiments, after the static computation graph is deployed, a static computation graph list is constructed, in which the input length of each deployed static computation graph and the inference time of each static computation graph are recorded.
[0138] The inference time of a static computation graph is the time from when input data enters the graph to when the execution results are generated. This inference time is determined by factors such as the complexity of the graph, the length of the input data, and the hardware performance required to run the computation. Inference time primarily depends on the hardware's computing power, memory bandwidth, storage speed, and degree of optimization.
[0139] In some exemplary embodiments of the present disclosure, the inference time of a static computation graph is determined offline. For example, each static computation graph is executed separately in a processor, and the inference time of each static computation graph is recorded. The inference time of each static computation graph is stored in a list.
[0140] In this embodiment, when deploying a static computation graph, we consider the impact of the underlying hardware capabilities on model execution efficiency to select a static computation graph with appropriate input length. We also determine the inference time based on hardware resources to improve the accuracy of the total inference time.
[0141] In some exemplary embodiments of the present disclosure, an offline method is used to determine the target static computation graph combination corresponding to the length of input data within a certain range. For each length of input data, a dynamic programming algorithm is used to determine the target static computation graph combination of that length, and the corresponding relationship between the length of the input data and the target static computation graph combination is recorded. For example: if the length range of the input data is 25-1000, then for any length between 25-1000, a dynamic programming algorithm is used to determine the target static computation graph combination of that length, and the relationship between the length and the target static computation graph combination is recorded. The target static computation graph combination includes at least one input length corresponding to a static computation graph. The target static computation graph combination is an optimal segmentation combination corresponding to the input data of that length, which may be the one with the shortest total inference time, the least memory usage, and so on.
[0142] During step S602, after obtaining the length of the input data, the input data length is used to query the previously recorded correspondence for the target static computation graph combination corresponding to the input data length. The input data is then divided into multiple sub-segments according to the input lengths of the static computation graphs in the target static computation graph combination. If the length of a sub-segment is less than the input length of the static computation graph, sub-segment 4 may be padded to meet the static computation graph input length requirement.
[0143] In step S606, the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment.
[0144] In step S608, the execution result of the static computation graph is output.
[0145] In this embodiment, an optimal segmentation combination is determined in an offline manner to improve the efficiency of online reasoning data.
[0146] like Figure 8 As shown, it is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, which includes the following steps.
[0147] In step S802 , input data is received.
[0148] In step S804, the input data is divided into multiple sub-segments based on the length of the input data and the input length of each static computation graph.
[0149] In step S806 , the last sub-segment among the multiple sub-segments is padded to match the input length of the corresponding static computation graph.
[0150] If the length of the last sub-segment after segmentation is less than the input length of the corresponding static computation graph, the last sub-segment is padded with preset elements so that the length of the last sub-segment reaches the input length of the static computation graph, and the padded last sub-segment is used as input to execute the corresponding static computation graph and generate the inference result of the input data, so as to realize inference on the sub-segment whose length is less than the input length of the static computation graph.
[0151] The preset element may be all 0s, all 1s, or any element in the sub-segment.
[0152] In some embodiments, the preset element can be filled in any position of the sub-segment, for example, at the beginning or end of the last sub-segment.
[0153] In some exemplary embodiments of the present disclosure, if the length of the last sub-segment after segmentation is smaller than the corresponding static computation graph, the static computation graph corresponding to the input length whose input length is greater than the length of the input data and whose difference between the two is the smallest is used as the static computation graph corresponding to the last sub-segment, and the last sub-segment is padded with preset elements so that the length of the last sub-segment reaches the input length of the static computation graph. The padded last sub-segment is used as input, and the corresponding static computation graph is executed to generate the inference result of the input data.
[0154] For example, if the input data length is 123, static computation graphs with input lengths greater than 123 include: static computation graph B with an input length of 128 and static computation graph C with an input length of 384. Using static computation graph C requires filling a large number of preset elements, while using static computation graph B with a small number of preset elements satisfies the input length of static computation graph B. Therefore, executing static computation graph B using the input data as input can reduce computing power waste.
[0155] In step S808 , the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment.
[0156] In step S810, the execution result of the static computation graph is output.
[0157] like Figure 9 As shown, it is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, which includes the following steps.
[0158] In step S902 , input data is received.
[0159] In step S904, a target static computation graph combination is determined based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph.
[0160] In some embodiments, after the static computation graph is deployed, a static computation graph list is constructed, in which the input length of each deployed static computation graph and the inference time of each static computation graph are recorded. During the execution of step S904, the input length of each static computation graph and the inference time of each static computation graph are directly read from the static computation graph list.
[0161] In some exemplary embodiments of the present disclosure, the input length of each static computation graph is determined according to business requirements, and the inference time of each static computation graph is determined by the hardware resources on which the static computation graph is deployed.
[0162] Different business scenarios have different requirements for input data length. In this embodiment, the input length of the static computation graph is not set arbitrarily, but is determined based on the specific business scenario and task objectives. For example, for a text classification task that needs to process long articles, a static computation graph with a longer input length can be selected; for a simple phrase analysis task that needs to process simple phrases or sentences, a static computation graph with a shorter input length can be selected.
[0163] In some exemplary embodiments of the present disclosure, the inference time of a static computation graph is determined offline. For example, each static computation graph is executed separately in a processor, and the inference time of each static computation graph is recorded. The inference time of each static computation graph is stored in a list.
[0164] In this embodiment, when deploying a static computation graph, we must not only consider the business needs for input data but also the impact of the underlying hardware capabilities on model execution efficiency to select a static computation graph with appropriate input length. We also determine inference time and switching time based on hardware resources to improve the accuracy of the total inference time.
[0165] In some exemplary embodiments of the present disclosure, various static computation graphs share model weights.
[0166] Model weights are learnable variables in a deep learning model, such as the weights and biases in a neural network. Models are optimized through the training process to enable the model to achieve a specific task. In a deep learning model, each static computation graph has an independent set of parameters. If multiple static computation graphs share model weights, this means that the graphs use the same weights and biases, rather than maintaining independent parameter sets.
[0167] When multiple static computation graphs share the same model parameters, it means that the same set of weights and biases will be used for calculation when executing multiple static computation graphs. This avoids the overhead of storing parameters separately for each static computation graph, thereby saving memory.
[0168] Based on the input data length, the input length of each static computation graph, and the inference time of each static computation graph, a dynamic programming algorithm is used to segment the input data into multiple sub-segments. The dynamic programming algorithm sets a state equation whose goal is to minimize the total inference time. The length of each sub-segment is one of the input lengths of the static computation graph. When selecting the sub-segment length, combinations close to the maximum input length of the static computation graph are prioritized to reduce the number of inferences and switching overhead. The total inference time includes the inference time of each static computation graph required to infer each sub-segment.
[0169] In some exemplary embodiments of the present disclosure, multiple groups of candidate static computation graphs are determined, and the number of sub-segments included in the input data corresponding to each group of candidate static computation graphs is determined, wherein the number of sub-segments included in the input data corresponding to each group of candidate static computation graphs is determined by the input length of each candidate static computation graph; the total inference time of each group of candidate static computation graphs is calculated, wherein the total inference time of each group of candidate static computation graphs is obtained by summing the inference time of each candidate static computation graph in the group.
[0170] Because multiple static computation graphs have been deployed in the model, they can work together to complete the inference task on the input data. If multiple static computation graphs work together to complete the inference task on the input data, then this combination can be considered a set of candidate static computation graphs.
[0171] A set of candidate static computation graphs may include multiple candidate static computation graphs. The multiple deployed static computation graphs are traversed, and all combinations whose total input length is greater than the total length of the input data are selected as multiple sets of candidate static computation graphs.
[0172] For example, static computation graphs A, B, and C with input lengths of 32, 128, and 384, respectively, are deployed in the model, and the total length of the input data is 500.
[0173] Candidate static computation graph combinations may include [static computation graph C and static computation graph B], [1 static computation graph C, 1 static computation graph B, and 1 static computation graph A], [5 static computation graphs B, 1 static computation graph A], and so on. In this embodiment, the candidate static computation graph combinations are merely illustrative and not limiting.
[0174] For each group of candidate static computation graphs, the input data is divided into multiple sub-segments according to the input length of each candidate static computation graph.
[0175] For example, for the static computation graph combination [static computation graph C and static computation graph B], the input data is segmented into sub-segment 1 with a length of 382 and sub-segment 2 with a length of 118. This segmentation method includes 2 sub-segments.
[0176] For example, for [4 static computation graphs B and 1 static computation graph A], the input data is segmented into sub-segment 1 with a length of 128, sub-segment 2 with a length of 128, sub-segment 3 with a length of 128, and sub-segment 4 with a length of 6. This segmentation method includes 4 sub-segments.
[0177] For example: for [16 static computation graphs A], the input data segmentation result includes that the number of sub-segments is 16, the length of the first 15 sub-segments is 32, and the length of the last sub-segment is 20.
[0178] Since each group of candidate static computation graphs includes multiple candidate static computation graphs, the inference time of each candidate static computation graph in the group is accumulated to obtain the total inference time of the group of candidate static computation graphs.
[0179] For example, for [static computation graph C and static computation graph B], the total inference time of this set of candidate static computation graphs is the inference time of static computation graph C plus the inference time of static computation graph B. Assume that the time for static computation graph C is 10ms. Then, the total inference time of this combination is 15ms.
[0180] For example, for [4 static computation graphs B and 1 static computation graph A], the total inference time for this set of candidate static computation graphs is the sum of the inference times of the 4 static computation graphs B and the 1 static computation graph A. Assume that the inference time for static computation graph B is 5ms and the time for static computation graph A is 2ms. Therefore, the total inference time for this combination is 18ms.
[0181] For example, for [16 static computation graphs A], the total inference time of this group of candidate static computation graphs is the sum of the inference times of the 16 static computation graphs B. If the time for static computation graph A is 2ms, the total inference time for this group is 32ms.
[0182] In step S906 , the input data is divided into a plurality of sub-segments based on the input length of each target static computation graph included in the target static computation graph combination.
[0183] The target static computation graph combination refers to a set of static computation graphs determined by the dynamic inference algorithm in step S804. The input length of each static computation graph is fixed. For example, static computation graph A can only process sub-segments with a length of 32, static computation graph B can process sub-segments with a length of 128, and static computation graph C supports sub-segments with a length of 384.
[0184] The input data is divided into a plurality of sub-segments according to the input lengths of the target static computation graphs included in the target static computation graph combination. The length of each sub-segment must meet the input requirements of at least one static computation graph.
[0185] For example, if the target static computation graph combination is [static computation graph C and static computation graph B] and the input data length is 500, then by splitting according to the input length of static computation graph C, we can obtain sub-segment 1: [t1, t2, ..., t384]; sub-segment 2: [t385, t386, ..., t500].
[0186] The length of the last sub-segment after segmentation is less than the input length of the static computation graph B. You can add a padding setting element to make the last sub-segment reach the input length of the static computation graph B.
[0187] In step S908, the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment;
[0188] In step S910, the execution result of the static computation graph is output.
[0189] In this embodiment, the dynamic programming algorithm can be used to find the optimal segmentation combination and improve the inference speed of the input data.
[0190] like Figure 10 As shown, it is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, which includes the following steps.
[0191] In step S1002 , input data is received.
[0192] Step S1002 provided in this embodiment is the same as the execution process of S802 provided in the above embodiment. For details, please refer to the description in the above embodiment. In this embodiment, no specific limitation is given.
[0193] In step S1004, a target static computation graph combination is determined based on the length of the input data, the input length of each static computation graph, the inference time of each static computation graph, and the switching time of each static computation graph.
[0194] Inference time refers to the time from when input data enters the static computation graph to when the execution results are generated. This time is determined by factors such as the complexity of the static computation graph, the length of the input data, and the hardware performance required to run the computation. Inference time primarily depends on the hardware's computing power, memory bandwidth, storage speed, and degree of optimization.
[0195] In some example embodiments of the present disclosure, the inference time of the static computation graph and the switching time between the static computation graphs are determined in an offline manner. For example, each static computation graph is executed in a processor, and the inference time of each static computation graph is recorded. The inference time of each static computation graph is stored in a list. For another example, each static computation graph is executed in a processor, and the switching time between the static computation graphs is recorded. The switching time between the static computation graphs is stored in the list of static computation graphs.
[0196] Based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph and the switching time between the static computation graphs, the input data is segmented by using a dynamic programming algorithm to obtain a plurality of subsegments, wherein the dynamic programming algorithm sets a state equation, and the objective of the state equation is to minimize the total inference time; the length of each subsegment is one of the input lengths of the static computation graphs, and when selecting the length of the subsegment, the combination close to the maximum input length of the static computation graph is preferred to reduce the inference times and the switching overhead. The total inference time includes the inference time of each static computation graph required to be used for inferring each subsegment and the switching time of each static computation graph required to be used for inferring each subsegment.
[0197] In some example embodiments of the present disclosure, a plurality of groups of candidate static computation graphs are determined, and the number of subsegments included in the input data corresponding to each group of candidate static computation graphs is determined, wherein the number of subsegments included in the input data corresponding to each group of candidate static computation graphs is determined by the input length of each candidate static computation graph in the group; and the total inference time of each group of candidate static computation graphs is calculated, wherein the total inference time of each group of candidate static computation graphs is obtained by accumulating the inference time of each candidate static computation graph in the group and the switching time of each static computation graph.
[0198] Since a plurality of static computation graphs have been deployed in the model, each static computation graph can cooperate with each other to complete the inference task of the input data. If the plurality of static computation graphs can cooperate with each other to complete the inference task of the input data, the combination can be used as a group of candidate static computation graphs.
[0199] A group of candidate static computation graphs can include a plurality of candidate static computation graphs. All combinations whose total input length is greater than the total length of the input data are used as a plurality of groups of candidate static computation graphs by traversing the plurality of static computation graphs that have been deployed.
[0200] For example, static computation graph A, static computation graph B and static computation graph C with input lengths of 32, 128 and 384 respectively are deployed in the model, and the total length of the input data is 500.
[0201] Candidate static computation graph combinations may include [static computation graph C and static computation graph B], [1 static computation graph C, 1 static computation graph B, and 1 static computation graph A], [5 static computation graphs B, 1 static computation graph A], and so on. In this embodiment, the candidate static computation graph combinations are merely illustrative and not limiting.
[0202] For each group of candidate static computation graphs, the input data is divided into multiple sub-segments according to the input length of each candidate static computation graph.
[0203] For example, for the static computation graph combination [static computation graph C and static computation graph B], the input data is segmented into sub-segment 1 with a length of 382 and sub-segment 2 with a length of 118. This segmentation method includes 2 sub-segments.
[0204] For example, for [4 static computation graphs B and 1 static computation graph A], the input data is segmented into sub-segment 1 with a length of 128, sub-segment 2 with a length of 128, sub-segment 3 with a length of 128, and sub-segment 4 with a length of 6. This segmentation method includes 4 sub-segments.
[0205] For example: for [16 static computation graphs A], the input data segmentation result includes that the number of sub-segments is 16, the length of the first 15 sub-segments is 32, and the length of the last sub-segment is 20.
[0206] Since each group of candidate static computation graphs includes multiple candidate static computation graphs, the inference time of each candidate static computation graph in the group and the switching time of each static computation graph are accumulated to obtain the total inference time of the group of candidate static computation graphs.
[0207] For example, for [static graph C and static graph B], the total inference time for this set of candidate static graphs is determined by the inference time of static graph C, the inference time of static graph B, and the switching time from static graph C to static graph B. Assuming that the inference time for static graph C is 10ms, the time for static graph B is 5ms, and the switching time from static graph C to static graph B is 1ms, the total inference time for this combination is 16ms.
[0208] For example, for [4 static graphs B and 1 static graph A], the total inference time for this set of candidate static graphs is the sum of the inference time of the 4 static graphs B, the inference time of the 1 static graph A, the switching time from the 3 static graphs B to the static graph B, and the switching time from the static graph B to the static graph A. Assuming the inference time for static graph B is 5ms, the time for static graph A is 2ms, the switching time from static graph B to static graph B is 1ms, and the switching time from static graph B to static graph A is 1ms. Therefore, the total inference time for this combination is 22ms.
[0209] For example, for [16 static computation graphs A], the total inference time for this group of candidate static computation graphs is the sum of the inference time of 16 static computation graphs B and the switching time of 15 static computation graphs A and static computation graph A. If the inference time of static computation graph A is 2ms and the switching time of static computation graph A is 1ms, the total inference time for this combination is 410ms.
[0210] In step S1006 , the input data is divided into a plurality of sub-segments based on the input length of each target static computation graph included in the target static computation graph combination.
[0211] In step S1008, the plurality of sub-segments are respectively used as input to execute a static computation graph corresponding to each sub-segment;
[0212] In step S1010, the execution result of the static computation graph is output.
[0213] In this embodiment, the inference speed of input data is improved by considering the influence of the switching time of the static computation graph on the total inference time in the dynamic programming algorithm.
[0214] In an exemplary embodiment, Figure 10 As shown in , according to the input length of each static computation graph and the inference time of each static computation graph, a dynamic programming algorithm is used to generate a target static computation graph combination, and then the input data is divided into multiple sub-segments according to the input length of the target static computation graph included in the target static computation graph combination. Figure 10As shown, 2 target static computations, i.e., static computation graph A and static computation graph B, are included in the target static computation graph combination; input data A is divided into subsegment 1 and subsegment 2. After the input data is divided into multiple subsegments, each static computation graph is executed. Subsegment 1 is taken as input to execute static computation graph A to obtain execution result a, and the execution result a is stored as context in the key-value cache, then subsegment 2 is taken as input to execute static computation graph B, and the execution result a is read from the key-value cache as context to assist understanding of subsegment 2, and finally execution result b is generated. Execution result b is taken as the inference result of input data and output. In addition, in processing in which multiple static computation graphs, such as static computation graph A, static computation graph B, static computation graph C, etc., are deployed, the model weights are shared between the static computation graphs.
[0215] In an example application scenario, Figure 12 is a schematic diagram of application of the data processing method according to an example embodiment of the present disclosure to a vehicle, as Figure 12 As shown, a dialogue page is displayed in the terminal interface. The dialogue page corresponds to the input operation of the user on the text box 121, obtains the text data input by the user, for example, obtains the input "today's weather is really good" by the user, and after receiving the text data input by the user, takes the text data as input data, executes the data inference method provided in the embodiment, can obtain the inference result of the input data, and displays it in the dialogue page. For example, the inference result "yes, today is sunny, it is a good day to go out for a walk!" is displayed in the text box 121 of the dialogue page. In the embodiment, only the application scenario of the data inference method is exemplarily described, but not limited.
[0216] In an example application scenario, Figure 13 is a schematic diagram of application of the data processing method according to an example embodiment of the present disclosure to a vehicle. As Figure 13 As shown, the vehicle 1300 includes a cabin 1310, which is assumed to include 4 seats (only for exemplarily description, the number of seats in the cabin is not limited by the present disclosure), for example, a main driver seat 1330, a co-driver seat 1340, a seat 1350 behind the main driver seat, and a seat 1360 behind the co-driver seat. The intelligent question and answer client 1320 (which can be abbreviated as client) is integrated in the vehicle 1300. The intelligent question and answer client 1320 can receive input data input by the user, and execute the data processing method provided in the above embodiment, and after obtaining the execution result, display the execution result on the interface of the intelligent question and answer client 1320, or play it in the form of voice.
[0217] Figure 14 is a block diagram of a data processing apparatus according to some embodiments of the present disclosure. Refer to Figure 14The device includes a receiving module 1410, a segmentation module 1420, an execution module 1430 and a generation module 1440.
[0218] Among them, the receiving module 1410 is used to receive input data; the segmentation module 1420 is used to divide the input data into multiple sub-segments; the execution module 1430 is used to use the multiple sub-segments as input to execute the static calculation graph corresponding to each sub-segment; and the output module 1440 is used to output the execution result of the static calculation graph.
[0219] In some exemplary embodiments of the present disclosure, the execution module 1430 is specifically configured to take the multiple sub-segments as input in sequence and execute the static computation graphs corresponding to the sub-segments in order, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
[0220] In some exemplary embodiments of the present disclosure, the segmentation module 1420 is specifically configured to segment the input data into multiple sub-segments according to the length of the input data.
[0221] In some exemplary embodiments of the present disclosure, the segmentation module 1420 is specifically configured to segment the input data into multiple sub-segments according to the same step size.
[0222] In some exemplary embodiments of the present disclosure, the segmentation module 1420 is specifically configured to divide the input data into a plurality of sub-segments according to the length of the input data, wherein the plurality of sub-segments include sub-segments of different lengths.
[0223] In some exemplary embodiments of the present disclosure, the segmentation module 1420 is specifically configured to divide the input data into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph.
[0224] In some exemplary embodiments of the present disclosure, the splitting module 1420: the combination determination unit 1510 is used to determine the target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph; the splitting unit 1520 is used to divide the input data into multiple sub-segments based on the input length of each static computation graph in the target static computation graph combination.
[0225] In some exemplary embodiments of the present disclosure, the combination determination unit 1421 is specifically used to determine the target static computation graph combination based on the length of the input data, the input length of each static computation graph, the inference time of each static computation graph, and the switching time of each static computation graph.
[0226] In some exemplary embodiments of the present disclosure, the inference time of each static computation graph is determined by deploying the static In some exemplary embodiments of the present disclosure, each static computation graph shares a model weight.
[0227] In some exemplary embodiments of the present disclosure, the apparatus further includes: a padding module 1610, configured to pad the last sub-segment of the multiple sub-segments to match an input length of a corresponding static computation graph.
[0228] Figure 17 1 is a block diagram illustrating a data processing electronic device 1700 according to some embodiments of the present disclosure. For example, the electronic device 1700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0229] Reference Figure 17 , the electronic device 1700 may include one or more of the following components: a processing component 1702 , a memory 1704 , a power component 1706 , a multimedia component 1708 , an audio component 1710 , an input / output (I / O) interface 1712 , a sensor component 1714 , and a communication component 1716 .
[0230] The processing component 1702 generally controls the overall operation of the electronic device 1700, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1702 may include one or more processors 1720 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1702 may include one or more modules to facilitate interaction between the processing component 1702 and other components. For example, the processing component 1702 may include a multimedia module to facilitate interaction between the multimedia component 1708 and the processing component 1702.
[0231] The memory 1704 is configured to store various types of data to support the operations of the device 1700. Examples of such data include instructions for any application or method operating on the electronic device 1700, contact data, phone book data, messages, pictures, videos, etc. The memory 1704 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0232] Power component 1706 provides power to various components of the electronic device 1700. The power component 1706 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 1700.
[0233] The multimedia component 1708 includes a screen to provide an output interface between the electronic device 1700 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 1708 includes a front camera and / or a rear camera. When the electronic device 1700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0234] The audio component 1710 is configured to output and / or input an audio signal. For example, the audio component 1710 includes a microphone (MIC) to receive an external audio signal when the electronic device 1700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1704 or transmitted via the communication component 1716. In some embodiments, the audio component 1710 also includes a speaker to output an audio signal.
[0235] The I / O interface 1712 provides an interface between the processing component 1702 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0236] Sensor assembly 1714 includes one or more sensors for providing various aspects of the status assessment of electronic device 1700. For example, sensor assembly 1714 can detect the open / closed state of electronic device 1700, the relative positioning of components, such as the display and keypad of electronic device 1700. Sensor assembly 1714 can also detect changes in the position of electronic device 1700 or a component of electronic device 1700, the presence or absence of user contact with electronic device 1700, the orientation or acceleration / deceleration of electronic device 1700, and changes in the temperature of electronic device 1700. Sensor assembly 1714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1714 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1714 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0237] The communication component 1716 is configured to facilitate wired or wireless communication between the electronic device 1700 and other devices. The electronic device 1700 can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, 5G, other communication standards, or a combination thereof. In some embodiments of the present disclosure, the communication component 1716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of the present disclosure, the communication component 1716 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0238] In some embodiments of the present disclosure, the electronic device 1700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0239] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 1704 including instructions, and the instructions can be executed by a processor 1720 of an electronic device 1700 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0240] A non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by the processor of a mobile terminal, enables the mobile terminal to perform a data processing method, the method comprising: receiving input data; dividing the input data into multiple sub-segments; and generating an execution result of the input data based on the multiple sub-segments and the static computation graph corresponding to each of the sub-segments.
[0241] Based on the same inventive concept, an embodiment of the present disclosure further provides a vehicle, which includes the electronic device as in the above embodiment.
[0242] Based on the same inventive concept, the present disclosure also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the data processing method of any one of the above-mentioned method embodiments. Since the principles for solving the problems in this computer program product embodiment are similar to those in the above-mentioned method embodiment, the implementation of this computer program product embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated here.
[0243] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0244] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0245] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0246] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A data processing method, characterized in that: include: Receive input data; dividing the input data into a plurality of sub-segments; Taking the multiple sub-segments as input, respectively, and executing a static computation graph corresponding to each sub-segment; Output the execution result of the static computation graph.
2. The data processing method according to claim 1, wherein: The executing of the static computation graph corresponding to each sub-segment using the multiple sub-segments as inputs respectively includes: The multiple sub-segments are sequentially used as input, and the static computation graphs corresponding to the sub-segments are executed in sequence, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
3. The data processing method according to claim 1, wherein: The step of dividing the input data into a plurality of sub-segments comprises: The input data is divided into a plurality of sub-segments according to the length of the input data.
4. The data processing method according to claim 3, wherein: The step of dividing the input data into a plurality of sub-segments according to the length of the input data comprises: The input data is divided into a plurality of sub-segments according to the same step size.
5. The data processing method according to claim 2, wherein: The step of dividing the input data into a plurality of sub-segments according to the length of the input data comprises: The input data is divided into a plurality of sub-segments according to the length of the input data, and the plurality of sub-segments include sub-segments of different lengths.
6. The data processing method according to any one of claims 3 to 5, characterized in that: The step of dividing the input data into a plurality of sub-segments according to the length of the input data comprises: The input data is divided into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph.
7. The data processing method according to claim 6, characterized in that: The step of dividing the input data into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph comprises: Determining a target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph; The input data is divided into a plurality of sub-segments based on the input length of each static computation graph in the target static computation graph combination.
8. The data processing method according to claim 7, characterized in that: The determining of a target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph includes: The target static computation graph combination is determined based on the length of the input data, the input length of each static computation graph, the inference time of each static computation graph, and the switching time of each static computation graph.
9. The data processing method according to claim 6, characterized in that: The inference time of each static computation graph is determined by the hardware resources on which the static computation graph is deployed.
10. The data processing method according to claim 6, characterized in that: The static computation graphs share model weights.
11. The data processing method according to claim 6, characterized in that: The method further comprises: The last sub-segment of the multiple sub-segments is padded to match the input length of the corresponding static computation graph.
12. A data processing device, characterized in that: The device comprises: A receiving module, configured to receive input data; A segmentation module, configured to divide the input data into a plurality of sub-segments; an execution module, configured to take the multiple sub-segments as input and execute a static computation graph corresponding to each sub-segment; An output module is used to output the execution result of the static calculation graph.
13. The data processing device according to claim 12, characterized in that The execution module is specifically configured to take the multiple sub-segments as input in sequence and execute the static computation graphs corresponding to the sub-segments in order, wherein the execution result of the previous static computation graph serves as the context of the next static computation graph.
14. The data processing device according to claim 12, characterized in that The segmentation module is specifically configured to divide the input data into a plurality of sub-segments according to the length of the input data.
15. The data processing device according to claim 14, characterized in that The segmentation module is specifically configured to divide the input data into a plurality of sub-segments based on the length of the input data and the input length of each static computation graph.
16. The data processing device according to claim 15, characterized in that The segmentation module is specifically used to determine the target static computation graph combination based on the length of the input data, the input length of each static computation graph, and the inference time of each static computation graph; and divide the input data into multiple sub-segments based on the input length of each static computation graph in the target static computation graph combination.
17. The data processing device according to claim 16, characterized in that The segmentation module is specifically used to determine the target static computation graph combination based on the length of the input data, the input length of each static computation graph, the inference time of each static computation graph, and the switching time of each static computation graph.
18. The data processing device according to claim 15, characterized in that The inference time of each static computation graph is determined by the hardware resources on which the static computation graph is deployed.
19. The data processing device according to claim 15, characterized in that: The static computation graphs share model weights.
20. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the steps of the data processing method according to any one of claims 1 to 11.
21. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to implement the steps of the data processing method according to any one of claims 1 to 11.
22. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, the steps of the data processing method according to any one of claims 1 to 11 are implemented.