Image data processing method and device, equipment, storage medium and program product
By segmenting the target tensor of the computer vision model at the highest dimension and distributing it to operators of multiple core groups, the problem of low inference performance of the computer vision model is solved, and more efficient parallel computing and performance improvement is achieved.
Patent Information
- Application Number
- CN202411882424.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
Computer vision models have low inference performance, and the existing technology is difficult to achieve effective parallel computing under the processing of complex computing graphs and multi-core group hardware architecture, resulting in limited performance improvement.
By segmenting the target tensor in the highest dimension, multiple sub-tensors are obtained and allocated to the operators of multiple core groups for processing, ensuring that the workload of each core group is balanced and avoiding waiting situations caused by uneven load.
It improves the speed of computer vision models for image data processing, improves the model's inference efficiency and performance, and ensures the consistency and integrity of target inference results.
Smart Images

Figure CN119940529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to an image data processing method, device, equipment, storage medium and program product. Background Art
[0002] With the continuous development of artificial intelligence technology, the types and number of computer vision models are increasing day by day, and the demand for accelerated computer vision model reasoning is also increasing. For the hardware architecture of multi-core groups, due to the high data transmission speed between core groups, it can be used to achieve parallel computing, thereby improving the reasoning speed of computer vision models. However, the more complex the structure of the computer vision model and the more types of operators, the more difficult it will be to apply parallel computing in practice, which will hinder the performance improvement of computer vision model reasoning.
[0003] In the related art, the performance of computer vision model reasoning is improved by the following two methods. The first method is to split the calculation graph, and the second method is to split the weights of the computer vision model. The first method is to put the operators used in the computer vision model reasoning on different core groups for calculation. Each core group only processes a sub-calculation graph of the computer vision model. After the calculation is completed, the result is transferred to another core group to continue processing the next sub-calculation graph of the computer vision model, thereby reducing the calculation time of each core group. When processing a small amount of input, this method can only be executed sequentially due to the dependency between the calculation graphs, and the performance of computer vision model reasoning cannot be improved. The second method is to split the weights contained in the operators in the computer vision model into different core groups. When calculating these operators, the results of all core group calculations need to be aggregated and a dimensionality reduction calculation is performed after completion to obtain the correct results, thereby reducing the calculation time of the operators on each core group. However, this method is more used in scenarios where the computer vision model is very large. The operations of aggregating the results and reducing the dimensionality usually bring additional computing overhead, resulting in reduced overall performance. Summary of the invention
[0004] In view of this, the present invention provides a method, apparatus, device, storage medium and program product for processing image data to solve the problem of low reasoning performance of computer vision models.
[0005] In a first aspect, the present invention provides a method for processing image data, comprising: obtaining a target tensor corresponding to the image data according to a received inference request, determining the highest dimension of the target tensor, wherein the image data is a numerical value used to describe image attributes; dividing the target tensor in the highest dimension to obtain multiple sub-tensors, and distributing the multiple sub-tensors to first operators corresponding to multiple core groups in a dividing order to obtain a distribution order, wherein the sub-tensors correspond one to one with the core groups; controlling a first operator to process the multiple sub-tensors to obtain multiple processing results, wherein the first operator is the first operator in an operator execution order corresponding to each core group to process the multiple sub-tensors, and the operator execution order is used to characterize the order in which the operators in each core group are executed; inputting the multiple processing results to a second operator corresponding to each core group, so that the second operator processes the multiple processing results in an operator execution order to obtain multiple target processing results, wherein the second operator is at least one operator executed after the first operator in the operator execution order; merging the multiple target processing results in the distribution order to obtain a target inference result.
[0006] According to the received inference request, the present invention obtains the target tensor corresponding to the image data, determines the highest dimension of the target tensor, and divides the target tensor in the highest dimension to obtain multiple sub-tensors. The target tensor of the present invention is stored in a continuous memory block when stored, and the data of the lower dimension is stored continuously in the memory. When the target tensor is divided in the highest dimension, the data of each sub-tensor in the lower dimension still maintains the original storage order. Therefore, the present invention divides the target tensor in the highest dimension to ensure the continuity of the target tensor. The present invention distributes multiple sub-tensors to the first operators corresponding to multiple core groups in the division order to obtain the distribution order. The sub-tensors correspond to the core groups one by one. The present invention evenly distributes multiple sub-tensors to multiple core groups, so that multiple core groups process multiple sub-tensors at the same time, thereby improving the processing speed of the computer vision model for the target tensor, thereby improving the performance of the computer vision model. The present invention controls the first operator to process multiple sub-tensors to obtain multiple processing results, and inputs the multiple processing results to the second operator corresponding to each core group, so that the second operator processes the multiple processing results in the operator execution order to obtain multiple target processing results, and merges the multiple target processing results in the allocation order to obtain the target reasoning result. The present invention controls multiple operators on each core group to process the target tensor, and merges the multiple target processing results in the allocation order to ensure the logical coherence and integrity of the target reasoning result. Compared with the related art, the present invention improves the speed of processing the target tensor, improves the efficiency of the computer vision model, and improves the performance of the computer vision model.
[0007] In an optional implementation, the target tensor is split in the highest dimension to obtain multiple sub-tensors, including: obtaining the array element corresponding to the highest dimension of the target tensor; obtaining the number of core groups, and splitting the array element corresponding to the highest dimension of the target tensor according to the number of core groups to obtain multiple sub-tensors, and the number of sub-tensors is consistent with the number of core groups.
[0008] The present invention divides the array elements corresponding to the highest dimension of the target tensor according to the number of core groups, so that the number of sub-tensors is consistent with the number of core groups, ensuring full utilization of computing resources. Since the workload of each core group is relatively balanced, the situation where some core groups wait for other core groups to complete tasks due to uneven load is avoided, thereby improving the reasoning efficiency of the computer vision model.
[0009] In an optional embodiment, after inputting multiple processing results into the second operator corresponding to each core group, the image data processing method also includes: determining whether there is missing data in the multiple processing results, and if there is missing data in the multiple processing results, obtaining missing information corresponding to the missing data; based on the missing information, querying the target data on the second core group, and copying the target data to the position corresponding to the missing data, the second core group is a core group other than the first core group where the missing data is located.
[0010] In the present invention, since the amount of data corresponding to the input area of the second operator on multiple core groups is fixed, when the processing result of the first operator of one of the core groups is missing, it is possible that during the calculation of the first operator, part of the data corresponding to the sub-tensor shifted to the first operator of other core groups. According to the missing information, a query is made on the second core group other than the first core group where the missing data is located to obtain the target data. The target data is the data that actually corresponds to the part. The target data is copied to the position corresponding to the missing data to ensure the integrity of the data and the accuracy of the target inference result obtained subsequently.
[0011] In an optional embodiment, determining whether there is missing data in multiple processing results includes: filling each processing result into the input area of the second operator, and determining whether there is a missing part in the input area; if there is a missing part in the input area, there is missing data in the processing result, and if there is no missing part in the input area, there is no missing data in the processing result.
[0012] In an optional embodiment, the missing information includes the missing number of missing data; based on the missing information, querying the target data on the second core group includes: determining whether there is overlapping data in the input area of the second operator on the second core group; if there is overlapping data in the input area of the second operator on the second core group, determining whether the number of overlapping data is consistent with the missing number; if the number of overlapping data is consistent with the missing number, taking the overlapping data as the target data.
[0013] In an optional implementation, before determining whether there is overlapping data in the input region of the second operator on the second core group, the method for processing image data further includes: calling a synchronization function to perform synchronization processing between the multiple core groups.
[0014] The present invention ensures synchronization between multiple core groups by calling a synchronization function, and when there is missing data in multiple processing results, ensures that the processing process on the second core group can support the query of target data.
[0015] In a second aspect, the present invention provides an image data processing device, including: a dimension determination module, used to obtain a target tensor corresponding to the image data according to a received inference request, and determine the highest dimension of the target tensor, wherein the image data is a numerical value used to describe image attributes; a sub-tensor allocation module, used to split the target tensor in the highest dimension to obtain multiple sub-tensors, and allocate the multiple sub-tensors to first operators corresponding to multiple core groups in a splitting order to obtain an allocation order, wherein the sub-tensors correspond to the core groups one by one; a first processing module, used to control the first operator to process the multiple sub-tensors to obtain multiple processing The first operator is the first operator that processes multiple sub-tensors in the operator execution order corresponding to each core group, and the operator execution order is used to characterize the order in which the operators in each core group are executed; the second processing module is used to input the multiple processing results to the second operator corresponding to each core group, so that the second operator processes the multiple processing results according to the operator execution order to obtain multiple target processing results, and the second operator is at least one operator executed after the first operator in the operator execution order; the reasoning result determination module is used to merge the multiple target processing results according to the allocation order to obtain the target reasoning result.
[0016] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the image data processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the image data processing method of the first aspect or any corresponding embodiment thereof.
[0018] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for processing image data of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 is a flowchart of a method for processing image data according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of a method for allocating data before segmentation according to an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of a method for allocating data after segmentation according to an embodiment of the present invention;
[0023] Figure 4 is a flow chart of another method for processing image data according to an embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of copying target data to a position corresponding to missing data according to an embodiment of the present invention;
[0025] Figure 6 is a flowchart of another method for processing image data according to an embodiment of the present invention;
[0026] Figure 7 is a structural block diagram of an image data processing device according to an embodiment of the present invention;
[0027] Figure 8 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0029] With the continuous development of artificial intelligence technology, the types and number of computer vision models are increasing day by day, and the demand for accelerated computer vision model reasoning is also increasing. The hardware architecture of the multi-core group of computer vision models can be used to achieve parallel computing due to the high data transmission speed between the core groups, thereby improving the performance of model reasoning; however, with the complexity of the computer vision model structure and the increase in the types of operators, parallel computing will be difficult to apply in practice, which will hinder the performance improvement of computer vision model reasoning.
[0030] In the related art, the performance of computer vision model reasoning is improved by the following two methods. The first method is to split the calculation graph, and the second method is to split the weights of the computer vision model. The first method is to put the operators used in the computer vision model reasoning on different core groups for calculation. Each core group only processes a sub-calculation graph of the computer vision model. After the calculation is completed, the result is transferred to another core group to continue processing the next sub-calculation graph of the computer vision model, thereby reducing the calculation time of each core group. When processing a small amount of input, this method can only be executed sequentially due to the dependency between the calculation graphs, and the performance of computer vision model reasoning cannot be improved. The second method is to split the weights contained in the operators in the computer vision model into different core groups. When calculating these operators, the results of all core group calculations need to be aggregated and a dimensionality reduction calculation is performed after completion to obtain the correct results, thereby reducing the calculation time of the operators on each core group. However, this method is more used in scenarios where the computer vision model is very large. The operations of aggregating the results and reducing the dimensionality usually bring additional computing overhead, resulting in reduced overall performance.
[0031] An embodiment of the present invention provides a method for processing image data, which divides a target tensor and distributes it to operators on multiple core groups for processing, so as to improve the processing speed of a computer vision model for image data.
[0032] According to an embodiment of the present invention, an embodiment of a method for processing image data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] In this embodiment, a method for processing image data is provided, which can be used in a computer device, and a computer vision model can be configured on the computer device. Figure 1 is a flowchart of a method for processing image data according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0034] Step S101, according to the received inference request, obtain the target tensor corresponding to the image data and determine the highest dimension of the target tensor.
[0035] The inference request is a request issued by a user to a computer vision model to infer or predict given image data. The inference request includes image data, which is a digital representation of an image. The image data contains various information about the image. The image data may include: pixel data, size and resolution data, depth data, and channel data. The target tensor is a value (array element) in the form of a multidimensional array corresponding to the image data. For example, for a color image with a resolution of 224*224 pixels, it can be represented by a three-dimensional tensor of (3, 224, 224), where 3 represents the number of color channels, and the two 224s represent the length and width of the image data, that is, the image data has 224 pixels in the vertical direction and 224 pixels in the horizontal direction. The dimension of the target tensor refers to the number of directions used to index elements in the target tensor. The highest dimension of the target tensor refers to the size of the first dimension in the shape representation of the target tensor. For example, for a three-dimensional tensor (3, 224, 224), the highest dimension is 3.
[0036] Step S102, split the target tensor in the highest dimension to obtain multiple sub-tensors, and distribute the multiple sub-tensors to the first operators corresponding to the multiple core groups in the splitting order to obtain the distribution order, and the sub-tensors correspond to the core groups one by one.
[0037] In some optional embodiments, the target tensor is split in the highest dimension to obtain multiple sub-tensors, including: obtaining the array element corresponding to the highest dimension of the target tensor; obtaining the number of core groups, and splitting the array element corresponding to the highest dimension of the target tensor according to the number of core groups to obtain multiple sub-tensors, and the number of sub-tensors is consistent with the number of core groups.
[0038] Among them, the storage of the target tensor in the computer memory is to arrange the array elements in a certain order. The array elements corresponding to the highest dimension of the target tensor are fixed. According to the number of core groups, the array elements corresponding to the highest dimension of the target tensor are evenly divided to obtain sub-tensors with the same number as the number of core groups, and the number of array elements in each sub-tensor is the same.
[0039] In some optional embodiments, the splitting order is the order in which the target tensor is split according to the arrangement order of the array elements. Exemplarily, the array elements corresponding to the highest dimension of the target tensor are split according to the arrangement order to obtain a first splitting result, a second splitting result, a third splitting result and a fourth splitting result. Then the splitting order is the first splitting result, the second splitting result, the third splitting result and the fourth splitting result.
[0040] In some optional embodiments, the allocation order refers to the order in which the multiple sub-tensors obtained by segmentation are allocated to the first operators corresponding to multiple core groups according to the segmentation order. For example, the first segmentation result is allocated to the first core group, the second segmentation result is allocated to the second core group, the third segmentation result is allocated to the third core group, and the fourth segmentation result is allocated to the fourth core group. The allocation order is the first core group, the second core group, the third core group, and the fourth core group.
[0041] The embodiment of the present invention divides the array elements corresponding to the highest dimension of the target tensor according to the number of core groups, so that the number of sub-tensors is consistent with the number of core groups, ensuring full utilization of computing resources. Since the workload of each core group is relatively balanced, the situation where some core groups wait for other core groups to complete tasks due to uneven load is avoided, thereby improving the reasoning efficiency of the computer vision model.
[0042] Step S103, controlling the first operator to process the multiple sub-tensors to obtain multiple processing results, wherein the first operator is the first operator to process the multiple sub-tensors in the operator execution order corresponding to each core group, and the operator execution order is used to represent the order in which the operators in each core group are executed.
[0043] Each core group includes multiple operators. The embodiment of the present invention divides multiple operators into two categories. The first category is the first operator, which is the first operator on the core group that processes the subtensor; the second category is the second operator, which is the operator on the core group other than the first operator. The second operator may include one operator or multiple operators. Exemplarily, when the second operator includes one operator, the operator execution order is the first operator, the second operator, wherein the output of the first operator is the input of the second operator.
[0044] Step S104, input the multiple processing results to the second operator corresponding to each core group, so that the second operator processes the multiple processing results according to the operator execution order to obtain multiple target processing results, and the second operator is at least one operator executed after the first operator in the operator execution order.
[0045] Among them, in the operator execution order, the output of the previous operator is the input of the next operator. When the second operator includes two operators, the operator execution order includes the execution order of the two operators. Exemplarily, if the two operators included in the second operator are operator 1 and operator 2, respectively, operator 1 is executed first and then operator 2 in the operator execution order, then the processing result is first input to operator 1 for processing to obtain the first processing result, and then the first processing result is input to operator 2 for processing to obtain the target processing result.
[0046] Step S105, merging multiple target processing results according to the allocation order to obtain a target inference result.
[0047] Among them, multiple target processing results are merged in the order of allocation. Since the multiple target processing results correspond to different parts after the original tensor is split, merging them in the correct order of allocation can completely summarize the information contained in each sub-part and restore the comprehensive information about the target tensor, thereby ensuring that the reasoning result is based on the complete input data, avoiding the reasoning deviation caused by the missing or incorrect combination of partial information, and ensuring the accuracy of the reasoning result to the greatest extent.
[0048] In some optional embodiments, such as Figure 2 As shown, this embodiment provides a pre-splitting data allocation method, including: inputting a target tensor to a first target operator on a core group so that the first target operator processes the target tensor to obtain a processing result, and inputting the processing result to a second target operator so that the second target operator processes the processing result to obtain a target inference result.
[0049] In the embodiment of the present invention, Figure 3 As shown, a method for allocating data after segmentation is provided, including: allocating multiple sub-tensors to a first operator on multiple core groups, so that the first operator processes the multiple sub-tensors to obtain multiple processing results, inputting the multiple processing results to a second operator, so that the second operator processes the multiple processing results to obtain multiple target processing results, and merging the multiple processing results to obtain a target inference result. Among them, the processing process of the first operator is the same as that of the first target operator, and the difference between the first operator and the first target operator is that the amount of data processed is different. The amount of data processed by the first target operator is the amount of data corresponding to the entire target tensor, and the amount of data processed by the first operator is the amount of data corresponding to the sub-tensors after the target tensor is segmented.
[0050] The image data processing method provided in this embodiment obtains the target tensor corresponding to the image data according to the received inference request, determines the highest dimension of the target tensor, and divides the target tensor in the highest dimension to obtain multiple sub-tensors. The target tensor of the embodiment of the present invention is stored in a continuous memory block, and the data of the lower dimension is stored continuously in the memory. When the target tensor is divided in the highest dimension, the data of each sub-tensor in the lower dimension still maintains the original storage order. Therefore, the embodiment of the present invention divides the target tensor in the highest dimension to ensure the continuity of the target tensor. The embodiment of the present invention distributes multiple sub-tensors to the first operators corresponding to multiple core groups according to the division order to obtain the distribution order. The sub-tensors correspond to the core groups one by one. The present invention evenly distributes multiple sub-tensors to multiple core groups, so that multiple core groups process multiple sub-tensors at the same time, thereby improving the processing speed of the computer vision model for the target tensor, thereby improving the performance of the computer vision model. The embodiment of the present invention controls the first operator to process multiple sub-tensors to obtain multiple processing results, and inputs the multiple processing results to the second operator corresponding to each core group, so that the second operator processes the multiple processing results in the operator execution order to obtain multiple target processing results, and merges the multiple target processing results in the allocation order to obtain the target reasoning result. The embodiment of the present invention controls multiple operators on each core group to process the target tensor, and merges the multiple target processing results in the allocation order to ensure the logical coherence and integrity of the target reasoning result. Compared with the related art, the embodiment of the present invention improves the speed of processing the target tensor, improves the efficiency of the computer vision model, and improves the performance of the computer vision model.
[0051] In this embodiment, a method for processing image data is provided, which can be used in a computer device. Figure 4 is a flowchart of another method for processing image data according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:
[0052] Step S401: According to the received inference request, obtain the target tensor corresponding to the image data and determine the highest dimension of the target tensor. The image data is a numerical value used to describe the image attributes. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0053] Step S402: Split the target tensor in the highest dimension to obtain multiple sub-tensors, and distribute the multiple sub-tensors to the first operators corresponding to the multiple core groups in the splitting order to obtain the distribution order. The sub-tensors correspond to the core groups one by one. For details, please refer to Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.
[0054] Step S403, control the first operator to process multiple sub-tensors to obtain multiple processing results. The first operator is the first operator to process multiple sub-tensors in the operator execution order corresponding to each core group. The operator execution order is used to represent the order in which the operators in each core group are executed. For details, please refer to Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0055] Step S404: Input the multiple processing results to the second operator corresponding to each core group. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0056] Step S405, determining whether there is missing data in the multiple processing results, and if there is missing data in the multiple processing results, obtaining missing information corresponding to the missing data.
[0057] In some optional embodiments, determining whether there is missing data in multiple processing results includes: filling each processing result into the input area of the second operator, and determining whether there is a missing part in the input area; if there is a missing part in the input area, there is missing data in the processing result, and if there is no missing part in the input area, there is no missing data in the processing result.
[0058] In some optional embodiments, the missing information includes the missing number of missing data; based on the missing information, querying the target data on the second core group includes: determining whether there is overlapping data in the input area of the second operator on the second core group; if there is overlapping data in the input area of the second operator on the second core group, determining whether the number of overlapping data is consistent with the missing number; if the number of overlapping data is consistent with the missing number, taking the overlapping data as the target data.
[0059] In some optional implementations, before determining whether there is overlapping data in the input region of the second operator on the second core group, the method further includes: calling a synchronization function to perform synchronization processing between the multiple core groups.
[0060] The embodiment of the present invention ensures synchronization between multiple core groups by calling a synchronization function, and when there is missing data in multiple processing results, ensures that the processing process on the second core group can support the query of target data.
[0061] Step S406 , according to the missing information, query the target data on the second core group, and copy the target data to the position corresponding to the missing data, the second core group being the core group other than the first core group where the missing data is located.
[0062] In some optional embodiments, such as Figure 5As shown, it is a schematic diagram of copying target data to the position corresponding to the missing data. There is a missing part in the input area of the second operator on the second core group. After querying multiple core groups, it is obtained that a part of the missing part is in the overlapping part of the processing result of the first operator on the first core group, and the other part of the missing part is in the overlapping part of the processing result of the first operator on the third core group. The overlapping part in the processing result of the first operator on the first core group and the overlapping part in the processing result of the first operator on the third core group are copied to the missing part of the input area of the second operator on the second core group to ensure the integrity of the data.
[0063] Step S407: Control the second operator to process the multiple processing results according to the operator execution order to obtain multiple target processing results. The second operator is at least one operator executed after the first operator in the operator execution order. For details, please refer to Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0064] Step S408: Combine multiple target processing results in the order of allocation to obtain a target inference result. Figure 1 Step S105 of the illustrated embodiment will not be described in detail here.
[0065] In the method for processing image data provided by the present embodiment, since the amount of data corresponding to the input area of the second operator on multiple core groups is fixed, when the processing result of the first operator of one of the core groups is missing, it is possible that during the calculation of the first operator, part of the data corresponding to the sub-tensor shifted to the first operator of other core groups. According to the missing information, a query is made on the second core group other than the first core group where the missing data is located to obtain the target data, which is the data that actually corresponds to the part. The target data is copied to the position corresponding to the missing data to ensure the integrity of the data and the accuracy of the target inference result obtained subsequently.
[0066] In this embodiment, a method for processing image data is provided, which can be used in a computer device. Figure 6 is a flowchart of another method for processing image data according to an embodiment of the present invention. Figure 6 As shown, the process includes the following steps:
[0067] Step S601, determine the core group corresponding to the current computer vision model, traverse each operator in the computer vision model, and determine the output data segmentation area of the core group corresponding to the operator.
[0068] Step S602 : According to the output data segmentation region of each operator on the corresponding core group in the computer vision model, reversely derive the input data segmentation region required by the operator on the corresponding core group.
[0069] In some optional embodiments, the parameters within the computer vision model are all determined. If the input shape of the computer vision model is determined at this time, the output shape of each operator can be deduced in sequence through the execution logic of the operator within the computer vision model until the final output shape of the computer vision model is also determined, thereby ensuring that the input and output ranges of multiple operators on multiple core groups in the computer vision model are known.
[0070] In some optional implementations, reverse deduction is used to determine which data in the input data segmentation region will affect the calculation result of the current output sub-region (there is a dependency in the calculation expression), and the range constituted by the input data is the derived input data segmentation region.
[0071] Step S603, according to the difference between different segmented areas on the same piece of data, the position offset relationship between the input area and the output area is calculated, and the position and size of the missing data in the input area are recorded.
[0072] In some optional embodiments, for the same set of data, there are two identities, the first is as the output of the previous operator, and the second is as the input of the next operator; when the sub-region derived when the data is used as the output is inconsistent with the sub-region derived when the data is used as the input, it is necessary to calculate the position offset between the sub-regions, and calculate the offset based on the position offset, and adjust the data based on the offset to ensure the integrity of the data.
[0073] Step S604, applying for storage space of the data segmentation area in the corresponding core group.
[0074] In some optional implementations, according to the data segmentation situation, corresponding storage space is applied for on each core group to ensure that the storage space can contain the union of the input and output sub-regions.
[0075] Step S605: After receiving the inference request, the input data is divided into different core groups, and each core group executes the calculation process of each operator in sequence.
[0076] The data in the embodiments of the present invention are all target tensors of image data.
[0077] Step S606: For areas with missing data, the current core group needs to copy the missing parts from other core groups containing missing data to the current core group. Before copying, synchronization between all core groups is required to ensure that the data is valid.
[0078] In the embodiment of the present invention, if there is missing data in the input data sub-region of the operator during the calculation process, it is necessary to copy the missing data from the corresponding other core groups to the current core group according to the location and size of the missing data. Before copying, a synchronization operation between all core groups is required to ensure that all data on the input data sub-region are ready.
[0079] In this embodiment, a device for processing image data is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0080] This embodiment provides a device for processing image data, such as Figure 7 As shown, including:
[0081] The dimension determination module 701 is used to obtain the target tensor corresponding to the image data according to the received inference request, and determine the highest dimension of the target tensor, where the image data is a numerical value used to describe image attributes.
[0082] The sub-tensor allocation module 702 is used to split the target tensor in the highest dimension to obtain multiple sub-tensors, and allocate the multiple sub-tensors to the first operators corresponding to multiple core groups in the splitting order to obtain the allocation order, and the sub-tensors correspond to the core groups one by one.
[0083] The first processing module 703 is used to control the first operator to process multiple sub-tensors to obtain multiple processing results. The first operator is the first operator to process multiple sub-tensors in the operator execution order corresponding to each core group. The operator execution order is used to represent the order in which the operators in each core group are executed.
[0084] The second processing module 704 is used to input multiple processing results to the second operator corresponding to each core group, so that the second operator processes the multiple processing results according to the operator execution order to obtain multiple target processing results. The second operator is at least one operator executed after the first operator in the operator execution order.
[0085] The inference result determination module 705 is used to merge multiple target processing results according to the allocation order to obtain a target inference result.
[0086] In some optional implementations, the sub-tensor allocation module 702 includes:
[0087] The data acquisition unit is used to obtain the array element corresponding to the highest dimension of the target tensor.
[0088] The slicing unit is used to obtain the number of core groups. According to the number of core groups, the array elements corresponding to the highest dimension of the target tensor are sliced to obtain multiple sub-tensors. The number of sub-tensors is consistent with the number of core groups.
[0089] In some optional implementations, the image data processing device further includes:
[0090] The first judgment module is used to judge whether there is missing data in the multiple processing results, and if there is missing data in the multiple processing results, obtain missing information corresponding to the missing data.
[0091] The data replication module is used to query the target data on the second core group according to the missing information and copy the target data to the position corresponding to the missing data. The second core group is a core group other than the first core group where the missing data is located.
[0092] In some optional implementations, the first determination module includes:
[0093] The second judgment unit is used to fill each processing result into the input area of the second operator, and judge whether there is a missing part in the input area. If there is a missing part in the input area, there is missing data in the processing result; if there is no missing part in the input area, there is no missing data in the processing result.
[0094] The second judgment unit is used to judge whether there is overlapping data in the input area of the second operator on the second core group.
[0095] The third judgment unit is used to judge whether the number of overlapping data is consistent with the number of missing data according to whether overlapping data exists in the input area of the second operator on the second core group.
[0096] The target data determination unit is used to use the overlapping data as target data if the number of overlapping data is consistent with the number of missing data.
[0097] In some optional implementations, the image data processing device further includes:
[0098] The synchronization processing module is used to call the synchronization function to perform synchronization processing between multiple core groups.
[0099] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0100] The image data processing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0101] The embodiment of the present invention also provides a computer device having the above Figure 7 The image data processing device is shown.
[0102] See also Figure 8 , Figure 8 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 8 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.
[0103] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0104] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0105] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0106] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0107] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0108] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0109] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0110] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for processing image data, characterized in that: The method comprises: According to the received inference request, obtain a target tensor corresponding to the image data, and determine the highest dimension of the target tensor; Splitting the target tensor in the highest dimension to obtain a plurality of sub-tensors, and allocating the plurality of sub-tensors to first operators corresponding to a plurality of core groups in a splitting order to obtain an allocation order, wherein the sub-tensors correspond to the core groups one by one; Controlling the first operator to process the multiple sub-tensors to obtain multiple processing results, wherein the first operator is the first operator to process the multiple sub-tensors in the operator execution order corresponding to each core group, and the operator execution order is used to represent the order in which the operators in each core group are executed; Inputting the multiple processing results to a second operator corresponding to each core group, so that the second operator processes the multiple processing results according to the operator execution order to obtain multiple target processing results, wherein the second operator is at least one operator executed after the first operator in the operator execution order; The multiple target processing results are merged according to the allocation order to obtain a target inference result.
2. The method according to claim 1, characterized in that The target tensor is split in the highest dimension to obtain multiple sub-tensors, including: Get the array element corresponding to the highest dimension of the target tensor; The number of core groups is obtained, and the array elements corresponding to the highest dimension of the target tensor are divided according to the number of core groups to obtain the multiple sub-tensors, where the number of sub-tensors is consistent with the number of core groups.
3. The method according to claim 1 or 2, characterized in that: After inputting the multiple processing results to the second operators corresponding to each core group, the method further includes: Determine whether there is missing data in the multiple processing results, and if there is missing data in the multiple processing results, obtain missing information corresponding to the missing data; According to the missing information, the target data is searched on the second core group, and the target data is copied to a position corresponding to the missing data, where the second core group is a core group other than the first core group where the missing data is located.
4. The method according to claim 3, characterized in that The determining whether there is missing data in the multiple processing results includes: Fill each processing result into the input area of the second operator, and determine whether there is a missing part in the input area; If there is a missing part in the input area, there will be missing data in the processing result; if there is no missing part in the input area, there will be no missing data in the processing result.
5. The method according to claim 3, characterized in that: The missing information includes the missing quantity of the missing data; and querying the target data on the second core group according to the missing information includes: Determining whether there is overlapping data in the input region of the second operator on the second core group; If there is overlapping data in the input region of the second operator on the second core group, determining whether the number of the overlapping data is consistent with the number of missing data; If the number of the overlapping data is consistent with the missing number, the overlapping data is used as the target data.
6. The method according to claim 5, characterized in that Before determining whether there is overlapping data in the input region of the second operator on the second core group, the method further includes: Call the synchronization function to synchronize multiple core groups.
7. An image data processing device, characterized in that: The device comprises: A dimension determination module, configured to obtain a target tensor corresponding to the image data according to the received inference request, and determine the highest dimension of the target tensor, wherein the image data is a numerical value used to describe an image attribute; A sub-tensor allocation module is used to split the target tensor in the highest dimension to obtain multiple sub-tensors, and allocate the multiple sub-tensors to the first operators corresponding to the multiple core groups according to the splitting order to obtain an allocation order, and the sub-tensors correspond to the core groups one by one; a first processing module, used to control the first operator to process the multiple sub-tensors to obtain multiple processing results, wherein the first operator is the first operator to process the multiple sub-tensors in the operator execution order corresponding to each core group, and the operator execution order is used to represent the order in which the operators in each core group are executed; a second processing module, configured to input the plurality of processing results to a second operator corresponding to each core group, so that the second operator processes the plurality of processing results according to the operator execution order to obtain a plurality of target processing results, wherein the second operator is at least one operator executed after the first operator in the operator execution order; The inference result determination module is used to merge the multiple target processing results according to the allocation order to obtain a target inference result.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image data processing method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the image data processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the image data processing method according to any one of claims 1 to 6.
Citation Information
Cited By
Task scheduling method, task scheduling device, electronic equipment and readable storage medium
CN121364936A