Data processing method and device, electronic equipment and storage medium

By dividing the sub-graph area into overlapping and non-overlapping areas and determining the area to be calculated according to the order of convolution operations, the problem of convolution operations in the prior art consumes a large amount of computing resources, and resource saving and calculation time reduction are achieved.

CN120219767APending Publication Date: 2025-06-27ARM TECH CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325562.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when the input feature map is large, it is necessary to divide it into multiple sub-graphs for convolutional operations, resulting in large amounts of data computing and consumption of a large amount of computing resources.

Method used

By dividing the area of ​​the sub-graph into overlapping areas and non-overlapping areas, and determining the area to be calculated according to the order of the sub-graph convolution operations, avoiding the sub-graph of the post-convolution operation repeatedly computeing the area overlapping with the sub-graph of the previous convolution operation, and using the storage space to cache the operation results.

Benefits of technology

It effectively saves computing resources, reduces computing time, and avoids repeated calculations of overlapping areas between sub-graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219767A_ABST
    Figure CN120219767A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of computers.The method comprises the steps that a plurality of sub-graphs formed by segmenting an input feature graph are obtained; wherein an overlapping region exists between the adjacent sub-graphs; determining a to-be-calculated region of the sub-graph according to the sequence of the convolution operation of the sub-graph; wherein the to-be-calculated area comprises a first area and a second area; convolution operation is carried out on the to-be-calculated areas of the sub-graphs in sequence, and operation results are cached in a storage space; wherein the operation result of the to-be-calculated region of the sub-graph and the operation result of the second region are respectively stored; and obtaining an operation result of the sub-graph from the storage space, and generating an output feature graph. According to the embodiment of the invention, in each layer of convolution operation, the data of the overlapping region of the adjacent sub-graphs is only calculated once, so that the technical problem that the overlapping region of the sub-graphs in the prior art is repeatedly calculated to cause the waste of calculation resources is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology. Specifically, this application relates to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of neural networks, in order to obtain better results, technicians have begun to design dedicated hardware devices such as processors for neural networks. Among them, the NPU (Neural Processing Unit) is a dedicated processor that can support multiple neural networks and can perform convolution calculations on input feature maps.

[0003] In the prior art, when the input feature map is large, the input feature map is often divided into multiple sub-maps for convolution operations, and finally the output data of the convolution operations of each sub-map is spliced together to generate the input feature map. However, the inventor found in the process of implementing the solution of this application that during the convolution operations of each sub-map, the amount of data operation is large, and the computing resources of the neural network processing unit need to be consumed. Therefore, how to reduce the problem of consuming a large amount of computing resources during the convolution operation has become an urgent task. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, electronic device, and storage medium to solve the technical problem of consuming a large amount of computing resources during the convolution operation in the prior art.

[0005] According to a first aspect of the embodiments of this application, a data processing method is provided. The method includes: Obtain a plurality of sub-maps obtained by dividing an input feature map; wherein, there is an overlapping area between adjacent sub-maps; Determine the area to be calculated of the sub-maps according to the order of the convolution operations of the sub-maps; wherein, the area to be calculated includes a first area and a second area; the first area includes the part that does not overlap with other sub-maps, and the second area includes the part that overlaps with other sub-maps; and the second area of the sub-map that is convolved first includes: the overlapping area that overlaps with the sub-map that is convolved later; and the second area of the sub-map that is convolved later does not include: the area that overlaps with the sub-map that is convolved first; Perform convolution operations on the area to be calculated of the sub-maps in sequence, and cache the operation results into the storage space; wherein, store the operation results of the area to be calculated of the sub-maps and the operation results of the second area respectively; Obtain the operation results of the sub-maps from the storage space and generate an output feature map.

[0006] As an alternative implementation, the convolution operation of the input feature map uses a preset convolution kernel and a stride; Before obtaining the multiple subgraphs obtained by splitting the input feature map, it further includes: According to the size of the preset convolution kernel and the stride, the input feature map is split into multiple subgraphs.

[0007] As an alternative implementation, the splitting of the input feature map into multiple subgraphs according to the size of the preset convolution kernel and the stride includes: According to the preset convolution kernel and the convolution stride, determine the size of the overlapping region when the convolution kernel slides in the corresponding convolution layer as the first size; Determine the size of the overlapping region of each subgraph according to the first size, and split the input feature map into multiple subgraphs according to the size of the overlapping region of each subgraph.

[0008] As an alternative implementation, before the subgraph stores the operation result of the region to be calculated and the operation result of the second region respectively, it further includes: Determine the coordinates of the second region in the subgraph, and obtain the operation result of the second region from the operation result of the region to be calculated in the subgraph according to the coordinates.

[0009] As an alternative implementation, the subgraph storing the operation result of the region to be calculated and the operation result of the second region respectively includes: For the subgraph with the non-empty second region, the operation result of the region to be calculated includes the operation result of the first region and the operation result of the second region; store the operation result of the region to be calculated and the operation result of the second region into their respective corresponding storage spaces; For the subgraph with the empty second region, the operation result of the region to be calculated includes the operation result of the first region, and store the operation result of the first region into the corresponding storage space.

[0010] As an alternative implementation, when the number of convolution operation layers in the region to be calculated of the subgraph is one layer; The generating of the output feature map by obtaining the operation result of the subgraph from the storage space includes: Read the operation result of the region to be calculated of the subgraph from the storage space, and splice the operation results of the region to be calculated of the subgraph to generate the output feature map.

[0011] As an alternative implementation, when the number of convolution operation layers in the region to be calculated of the subgraph is multiple layers; Obtaining the operation result of the sub-graph from the storage space and generating an output feature map includes: For non-last-layer convolution operations, the sub-graph of the prior convolution operation obtains a first operation result from the storage space, and uses the first operation result as the input data for the next convolution layer of the sub-graph of the prior convolution operation; wherein, the first operation result is the operation result of the area to be calculated in the current convolution layer of the sub-graph of the prior convolution operation; the sub-graph of the subsequent convolution operation obtains a second operation result and a third operation result from the storage space; wherein, the second operation result is the operation result of the area to be calculated in the current convolution layer of the sub-graph of the subsequent convolution operation, and the third operation result is the operation result of the second area in the current convolution layer of the sub-graph of the prior convolution operation, and the second operation result and the third operation result are concatenated as the input data for the next convolution layer of the sub-graph of the subsequent convolution operation; For the last-layer convolution operation, the operation result of the area to be calculated of the sub-graph of the prior convolution operation and the operation result to be calculated of the sub-graph of the subsequent convolution operation are concatenated to generate an output feature map.

[0012] According to the second aspect of the embodiments of the present application, a data processing device is provided, and the device includes: A first processing module, configured to obtain multiple sub-graphs obtained by splitting an input feature map; wherein, there is an overlapping area between adjacent sub-graphs; A second processing module, configured to determine the area to be calculated of the sub-graph according to the order of convolution operations of the sub-graph; wherein, the area to be calculated includes a first area and a second area; the first area includes the part that does not overlap with other sub-graphs, and the second area includes the part that overlaps with other sub-graphs; and the second area of the sub-graph of the prior convolution operation includes: the overlapping area that overlaps with the sub-graph of the subsequent convolution operation; and the second area of the sub-graph of the subsequent convolution operation does not include: the area that overlaps with the sub-graph of the prior convolution operation; A third processing module, configured to sequentially perform convolution operations on the area to be calculated of the sub-graph and cache the calculation results in the storage space; wherein, the calculation results of the first area and the calculation results of the second area are stored separately; A fourth processing module, configured to obtain the output data of the sub-graph from the storage space and generate the output feature map.

[0013] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method described in any item of the first aspect.

[0014] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspect are implemented.

[0015] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are as follows: In the embodiments of the present application, the area of the sub-graph is divided into overlapping areas and non-overlapping areas, and according to the order of sub-graph convolution operations, the area to be calculated of the sub-graph in the subsequent convolution operation does not include the area overlapping with the sub-graph in the previous convolution operation. In this way, the sub-graph in the subsequent convolution operation does not need to calculate the area overlapping with the sub-graph in the previous convolution operation; at the same time, the sub-graph in the previous convolution operation stores the operation result of the overlapping area in the storage space, so that the sub-graph in the subsequent convolution operation can directly obtain the operation result of the overlapping area from the storage space, ensuring that the overlapping area between sub-graphs will not be repeatedly calculated, saving computing resources and reducing computing time. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.

[0017] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present application; Figure 2a It is a schematic diagram of an input feature map including two sub-graphs provided by an embodiment of the present application; Figure 2b It is a schematic flowchart of a convolution operation on an input feature map including two sub-graphs in a related art provided by an embodiment of the present application; Figure 2c It is a schematic flowchart of a convolution operation on two sub-graphs provided by an embodiment of the present application; Figure 3 It is a schematic diagram of dividing an input feature map in a related art provided by an embodiment of the present application; Figure 4 It is a schematic diagram of dividing an input feature map provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0018] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0019] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0020] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0021] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0022] With the continuous development of neural networks, in order to obtain better results, technicians have begun to design dedicated hardware devices such as processors for neural networks. Among them, the NPU (Neural Processing Unit) is a dedicated processor that can support multiple neural networks and can perform convolution calculations on the input feature map.

[0023] In the prior art, when the input feature map is large, the input feature map is often divided into multiple sub-maps for convolution operations, and finally the output data of the convolution operations of each sub-map are stitched together to generate an output feature map. However, the inventor found during the implementation of the solution of this application that the size of the divided sub-maps needs to match the convolution kernel and the convolution stride of the convolution operation; therefore, when dividing the input feature map, adjacent sub-maps need to include overlapping regions, and each sub-map will calculate the overlapping regions during the convolution operation, resulting in repeated calculations of the overlapping regions, wasting the computing resources of the neural network processing unit. At the same time, it also increases the system power consumption. Therefore, it has become an urgent task to solve the problem that the overlapping regions of each sub-map in the prior art are repeatedly calculated, wasting computing resources and increasing system power consumption.

[0024] The data processing method, device, electronic device and storage medium provided by this application are intended to solve the above technical problems in the prior art.

[0025] The technical solution of the embodiment of this application and the technical effects produced by the technical solution of this application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can be referenced, learned from or combined with each other. For the same terms, similar features and similar implementation steps in different embodiments, they will not be described repeatedly.

[0026] It can be understood that in the data processing method provided by the embodiments of the present disclosure, any method step can be executed by an electronic device and / or a server. All steps in the method can be independently executed by the electronic device or the server, or jointly executed by the electronic device and the server.

[0027] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The electronic device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto.

[0028] Subsequently, the server is used as the execution subject to introduce the embodiments of the present disclosure. However, this does not constitute a limitation to the embodiments of the present disclosure.

[0029] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of this application. As shown in the figure, the method includes: S101. Obtain multiple sub-maps obtained by dividing an input feature map; wherein, there are overlapping regions between adjacent sub-maps.

[0030] In the embodiments of the present application, the input feature map generally refers to the intermediate feature representation in a deep learning model (such as a convolutional neural network). To process large-sized input feature maps, they are usually divided into multiple smaller sub-maps to facilitate parallel processing or reduce computational complexity.

[0031] Specifically, when dividing the input feature map, the size of the convolutional kernel and the convolutional stride need to be considered so that the size of the divided sub-maps matches the size of the convolutional kernel and the convolutional stride; therefore, a certain overlapping area is reserved between adjacent sub-maps. The purpose of this overlapping area is to ensure that boundary information is not lost during subsequent feature extraction or processing, thereby maintaining the continuity of the feature map.

[0032] S102. Determine the area to be calculated for the sub-map according to the order of sub-map convolution operations; wherein, the area to be calculated includes a first area and a second area; the first area includes the part that does not overlap with other sub-maps, and the second area includes the part that overlaps with other sub-maps; and the second area of the sub-map in the prior convolution operation includes: the overlapping area that overlaps with the sub-map in the subsequent convolution operation; and the second area of the sub-map in the subsequent convolution operation does not include: the area that overlaps with the sub-map in the prior convolution operation.

[0033] In the embodiments of the present application, first, the order of convolution operations for each sub-map needs to be determined. This can be determined according to specific algorithm designs or application requirements. For example, it can be sorted according to the position order, size order, or other specific rules of the sub-maps in the input feature map. Further, for each sub-map, according to its position and the overlapping situation with adjacent sub-maps, it is divided into two areas: a first area and a second area; wherein, the first area includes all pixels or feature points in the sub-map that do not overlap with other sub-maps. When performing convolution operations, the processing of this part of the area is relatively independent and does not need to consider the interaction with other sub-maps; the second area includes all pixels or feature points in the sub-map that overlap with other sub-maps.

[0034] Specifically, for the sub-map in the prior convolution operation, its second area should include all areas that overlap with the sub-map in the subsequent convolution operation; for the sub-map in the subsequent convolution operation, its second area does not include the area that overlaps with the sub-map in the prior convolution operation; this is to ensure that the sub-map in the subsequent convolution operation can utilize the results of the previous sub-map when processing the overlapping area, thereby avoiding repeated calculations and maintaining feature consistency.

[0035] Specifically, in the embodiments of the present application, the region of the sub-graph is divided into a non-overlapping region and an overlapping region. The non-overlapping region is the region independently included in the sub-graph. Therefore, there is no duplicate calculation in the non-overlapping region. For the overlapping region, if the sub-graph in the prior convolution operation includes the overlapping region, then the sub-graph in the subsequent convolution operation needs to remove this overlapping region. In this way, the regions to be calculated for each sub-graph do not include overlapping regions, thereby ensuring that the overlapping region will not be repeatedly calculated during each layer of convolution operation, wasting the computing resources of the neural network processor.

[0036] It should also be noted that if the overlapping region of the sub-graph in the subsequent convolution operation includes regions overlapping with multiple sub-graphs, then the second region of the sub-graph in the subsequent convolution operation only removes the region overlapping with the prior convolution sub-graph. Therefore, the second region of the sub-graph in the subsequent convolution operation is not necessarily empty. The following is an example for illustration. The input feature map is divided into three sub-graphs, namely sub- Figure 1 graph 1, sub-graph 2 and sub- Figure 3 graph 3; among them, sub- Figure 1 graph 1 includes a non-overlapping region a and a region b overlapping with sub-graph 2; sub-graph 2 includes a non-overlapping region c, a region b overlapping with sub- Figure 1 graph 1, and a region d overlapping with sub- Figure 3 graph 3; sub- Figure 3 graph 3 includes a non-overlapping region e and a region d overlapping with sub-graph 2. The convolution operation sequence is sub- Figure 1 graph 1, sub-graph 2 and sub- Figure 3 graph 3 in sequence. Then, the first region of sub- Figure 1 graph 1 is a, and the second region is b; the first region of sub-graph 2 is c, and the second region is d; the first region of sub- Figure 3 graph 3 is e, and the second region is empty; for sub-graph 2, sub- Figure 1 graph 1 is the prior convolution operation sub-graph, which already includes the overlapping region d, but does not include the region d where sub-graph 2 overlaps with sub- Figure 3 graph 3. Therefore, the second region of sub-graph 2 is not empty; for sub- Figure 3 graph 3, sub-graph 2 is the prior convolution operation sub-graph, and the second region of sub-graph 2 already includes the overlapping region d. Therefore, the second region of sub- Figure 3 graph 3 does not include the region d overlapping with sub-graph 2, that is: the second region of sub- Figure 3 graph 3 is empty.

[0037] S103. Perform convolution operations on the regions to be calculated of the sub-graphs in sequence, and cache the operation results into the storage space; wherein, store the operation results of the regions to be calculated of the sub-graphs and the operation results of the second regions respectively.

[0038] In the embodiments of the present application, according to the order of sub - graph convolution operations, convolution operations are sequentially performed on the regions to be calculated in the sub - graph, and the operation results of the regions to be calculated in the sub - graph and the operation results of the second region are respectively stored in the storage space.

[0039] Specifically, in the embodiments of the present application, the regions to be calculated in the sub - graph include a first region and a second region. After performing convolution operations on the sub - graph, the operation results of the regions to be calculated are obtained; in other words, the operation results of the regions to be calculated include the operation results of the first region and the operation results of the second region; further, the operation results of the regions to be calculated in the sub - graph and the operation results of the second region are respectively stored in the storage space.

[0040] S104. Obtain the operation results of the sub - graph from the storage space and generate an output feature map.

[0041] Specifically, in the embodiments of the present application, the sub - graph that performs the prior convolution operation stores the operation results of the overlapping region with the sub - graph that performs the subsequent convolution operation in the storage space. In this way, the sub - graph that performs the subsequent convolution operation can obtain the operation results of the overlapping region with the sub - graph that performs the prior convolution operation from the storage space. The sub - graph that performs the subsequent convolution operation concatenates the operation results of the current convolution layer and the operation results of the overlapping region as the input data for the next convolution layer; for the last convolution layer, the sub - graph that performs the subsequent convolution operation concatenates the operation results of the current convolution layer and the operation results of the overlapping region as the operation results of the subsequent sub - graph, and concatenates the operation results of all sub - graphs to generate an output feature map.

[0042] In the embodiments of the present application, the region of the sub - graph is divided into an overlapping region and a non - overlapping region, and according to the order of sub - graph convolution operations, the regions to be calculated in the sub - graph that performs the subsequent convolution operation do not include the regions overlapping with the sub - graph that performs the prior convolution operation. In this way, the sub - graph that performs the subsequent convolution operation does not calculate the regions overlapping with the sub - graph that performs the prior convolution operation; at the same time, the sub - graph that performs the prior convolution operation stores the operation results of the overlapping region in the storage space, so that the sub - graph that performs the subsequent convolution operation can directly obtain the operation results of the overlapping region from the storage space, ensuring that the overlapping regions between sub - graphs will not be calculated repeatedly, saving computing resources and reducing computing time.

[0043] To facilitate those skilled in the art to more clearly understand the technical solution of the present application, the following will be further illustrated by several embodiments.

[0044] Figure 2aSchematic diagram of an input feature map including two sub - graphs provided by an embodiment of the present application; as shown in the figure, the input feature map is divided into sub - graph a and sub - graph b; among them, there is an overlapping area o between sub - graph a and sub - graph b, sub - graph a also includes a non - overlapping area m, and sub - graph b also includes a non - overlapping area n. After the input feature map is divided into sub - graphs, convolution operations need to be performed on sub - graph a and sub - graph b respectively, and finally, the operation results of the convolution operation of sub - graph a and the result of the convolution operation of sub - graph b are spliced together to generate an output feature map.

[0045] Figure 2b Flow chart of a related technology for performing convolution operation on an input feature map including two sub - graphs provided by an embodiment of the present application; as Figure 2b shown, sub - graph a and sub - graph b need to perform four - layer convolution operations, and each layer of convolution operation has a corresponding convolution operator; for the first - layer convolution operation, the area to be calculated for sub - graph a is m + o, and the area to be calculated for sub - graph b is n + o. Perform convolution operation on the area to be calculated for sub - graph a to obtain m1 + o1, and perform convolution operation on the area to be calculated for sub - graph b to obtain n1 + o1; where m1 is the result of the convolution operation of the non - overlapping area m of sub - graph a, o1 is the result of the convolution operation of the overlapping area o between sub - graph a and sub - graph b, and n1 is the result of the convolution operation of the non - overlapping area of sub - graph b.

[0046] Specifically, when sub - graph a and sub - graph b perform convolution operations, the overlapping area o is convolved twice; further, during the second - layer convolution, third - layer convolution, and fourth - layer convolution operations, the results of the convolution operations of the overlapping area o of sub - graph a and sub - graph b are o2, o3, and o4 respectively; that is to say, in each layer of convolution operation, sub - graph a and sub - graph b will calculate the overlapping area, resulting in the overlapping area being calculated repeatedly, wasting the computing resources of the neural network processor, and at the same time, increasing the computing time.

[0047] Figure 2c Flow chart of performing convolution operation on an input feature map including two sub - graphs provided by an embodiment of the present application, as Figure 2c shown, the first area of sub - graph a is m, the second area is o, and the area to be calculated is m + o; the first area of sub - graph b is n, the second area is empty, and the area to be calculated is n; sub - graph a and sub - graph b need to perform four - layer convolution operations, and each layer of convolution operation has a corresponding convolution operator; for the first - layer convolution operation, the operation result of the convolution operation of the area to be calculated for sub - graph a is m1 + o1, and the result of the convolution operation of the area to be calculated for sub - graph b is n1; where m1 is the result of the convolution operation of the non - overlapping area m of sub - graph a, o1 is the result of the convolution operation of the overlapping area o between sub - graph a and sub - graph b, and n1 is the result of the convolution operation of the non - overlapping area of sub - graph b.

[0048] Before performing the second-layer convolution operation, sub-graph a and sub-graph b respectively store the operation results of the first-layer convolution operation in the storage space. At the same time, sub-graph a also stores the operation results of the second region in the storage space. In this way, when sub-graph b performs the second-layer convolution operation, it can obtain the operation result o1 of the overlapping region o of the first-layer convolution operation from the storage space, and splice n1 and o1 together as the input data for the second-layer convolution operation of sub-graph b; by analogy, when sub-graph b performs the third-layer convolution operation, it obtains the operation result o2 of the overlapping region of the second-layer convolution operation from the storage space, and splices n2 and o2 together as the input data for the third-layer convolution operation; when sub-graph b performs the fourth-layer convolution operation, it obtains the operation result o3 of the overlapping region of the third-layer convolution operation from the storage space, and splices n3 and o3 together as the input data for the fourth-layer convolution operation; finally, the operation result m4 + o4 of the fourth-layer convolution operation of sub-graph a and the output result n4 of the fourth-layer convolution operation of sub-graph b are spliced together to generate the output feature map.

[0049] Specifically, in the embodiment of the present application, sub-graph a is the sub-graph of the prior convolution operation, and the operation result of sub-graph a already includes the operation result of the overlapping region o. Sub-graph b obtains the operation result of the overlapping region of the corresponding convolution layer from the storage space; it can be seen that the technical solution provided by the embodiment of the present application does not need to perform repeated calculations on the overlapping region, reduces the data volume of the convolution operation, and saves computing resources.

[0050] Based on the above embodiments, as an optional embodiment, the convolution operation of the input feature map uses a preset convolution kernel and stride. Before obtaining the multiple sub-graphs obtained by splitting the input feature map, it further includes: According to the size of the preset convolution kernel and the stride, the input feature map is split into multiple sub-graphs.

[0051] Specifically, the convolution kernel size defines the width and height of the convolution kernel, and the convolution stride defines the number of pixels that the convolution kernel moves each time on the input feature map; when the convolution stride is smaller than the convolution kernel size, an overlapping region will be generated during the sliding of the convolution kernel, and the size of the overlapping region depends on the relative relationship between the convolution kernel size and the stride. In the case of no padding, the size of the sub-graph can be calculated from the size of the input feature map, the convolution kernel size, and the stride.

[0052] It should be noted that in the embodiment of the present application, when the convolution operation includes multiple layers, the convolution kernel and stride used in each layer of convolution operation can be the same or different, and need to be determined according to the actual situation.

[0053] In the embodiments of the present application, for the input feature map, the convolution kernels and strides used in the convolution operation are preset, and the input feature map is segmented by the preset convolution kernel size and stride to ensure that the size of each sub-map matches the preset convolution kernel and stride; at the same time, the size of the overlapping area of the sub-maps can also be reduced, effectively reducing the data volume of the convolution operation.

[0054] Based on the above embodiments, as an alternative embodiment, the input feature map is segmented into multiple sub-maps according to the size of the preset convolution kernel and the stride, including: Determine the size of the overlapping area when the convolution kernel slides in the corresponding convolution layer according to the preset convolution kernel and the convolution stride, as the first size; Determine the size of the overlapping area of each sub-map according to the first size, and segment the input feature map into multiple sub-maps according to the size of the overlapping area of each sub-map.

[0055] Specifically, when the convolution stride is less than the convolution kernel size, an overlapping area will be generated during the sliding of the convolution kernel. The size of the overlapping area depends on the relationship between the convolution kernel size and the stride. In the related art, usually, the size of the overlapping area of the last convolution operation is first determined, and then the sizes of the overlapping areas of the previous convolution operations are sequentially deduced backward. The feature map is split according to the size of the overlapping area of each convolution operation. Although this can ensure that the size of the overlapping area of the sub-maps matches the convolution kernel size and the steps during each convolution operation, the size of the overlapping area of the sub-maps of the previous convolution operation needs to consider the size of the overlapping area of the sub-maps of the next convolution operation, which will lead to a linear increase in the size of the overlapping area; when the number of convolution layers is large, this situation will become more obvious.

[0056] Specifically, in the embodiments of the present application, the parameters of the convolution operation can be preset in a software configuration manner, including but not limited to: the number of convolution layers, the convolution kernel size, the stride, and the number of sub-maps of each convolution operation. According to the parameters configured by the software, the input feature map is segmented to obtain each sub-map.

[0057] Optionally, in the embodiments of the present application, the convolution kernel size and stride of each convolution operation can be the same or different, and need to be determined according to the actual situation.

[0058] In the embodiments of the present application, the size of the overlapping area is determined according to the convolution kernel size and the stride, and the input feature map is segmented according to the size of the overlapping area; in other words, in the embodiments of the present application, when the input feature map is segmented into sub-maps, only the convolution kernel size and stride of the current convolution layer need to be considered, and the size of the overlapping area is determined according to the convolution kernel size and stride of the current convolution layer, without considering the size of the overlapping area during the next convolution operation, which can effectively reduce the size of the overlapping area, and thus effectively reduce the overlapping data.

[0059] To facilitate a clearer understanding by those skilled in the art of the method for segmenting an input feature map in the related art and the difference from the method for segmenting an input feature map in the embodiments of the present application, it will be illustrated below by examples.

[0060] Please refer to Figure 3 , which exemplarily shows a schematic diagram of segmenting an input feature map in a related art provided by the embodiments of the present application. As Figure 3 shown, for a square region of the input feature map, three layers of convolution operations need to be performed. The convolution kernel size for each layer of convolution operation is , and the stride is ; according to the convolution kernel size and the stride, it can be known that the overlapping region size of the third layer of convolution operation should be . By backward deduction in sequence, the overlapping region size of the second layer of convolution operation should be , and the overlapping region size of the first layer of convolution operation is . By determining the overlapping region size during each layer of convolution operation, the input feature map is segmented; as Figure 3 shown, the first figure is the input feature map. After determining that the overlapping region size of the first layer of convolution operation is , the input feature map is segmented into 4 sub - maps, namely the first sub - map, the second sub - map, the third sub - map, and the fourth sub - map; among them, the first sub - map includes regions 1, 2, 4, and 5; the second sub - map includes regions 2, 3, 5, and 6; the third sub - map includes regions 4, 5, 7, and 8; the fourth sub - map includes regions 5, 6, 8, and 9; among them, region 5 is the overlapping region of the four sub - maps, with a size of ; after the first layer of convolution operation, the second figure is obtained. The region numbers included in the first sub - map to the fourth sub - map do not change, and region 5 is the overlapping region of the four sub - maps, with a size of ; after the second layer of convolution operation, the third figure is obtained. The region numbers included in the first sub - map to the fourth sub - map do not change, and region 5 is the overlapping region of the four sub - maps, with a size of ; after the third layer of convolution operation, the fourth figure is obtained, that is, the output feature map; among them, the first sub - map includes region 1, the second sub - map includes region 2, the third sub - map includes region 3, the fourth sub - map includes region 4, and there is no overlapping region between adjacent sub - maps.

[0061] Please refer to Figure 4 , which exemplarily shows a schematic diagram of segmenting an input feature map provided by the embodiments of the present application. As Figure 4 shown, for a square region of the input feature map, three layers of convolution operations need to be performed. The convolution kernel size for each layer of convolution operation is , and the stride is ; According to the convolution kernel size and the stride, the overlapping region size of the third - layer convolution operation should be , therefore, the overlapping region size during each layer of convolution operation is . After determining the overlapping region size during each layer of convolution operation, the input feature map can be segmented. The first figure is the input feature map. After determining that the overlapping region size of the first - layer convolution operation is , the input feature map is segmented into 4 sub - maps, namely the first sub - map, the second sub - map, the third sub - map, and the fourth sub - map; among them, the first sub - map includes regions 1, 2, 4, and 5; the second sub - map includes regions 2, 3, 5, and 6; the third sub - map includes regions 4, 5, 7, and 8; the fourth sub - map includes regions 5, 6, 8, and 9; as can be seen from the figure, the region sizes of the first sub - map to the fourth sub - map are not exactly the same. The region size of the first sub - map is , the region size of the second sub - map is , the region size of the third sub - map is , the region size of the fourth sub - map is ; among them, region 5 is the overlapping region of the four sub - maps, with a size of . The segmented sub - maps are successively subjected to the first - layer convolution operation and the second - layer convolution operation, and the second figure and the third figure are obtained respectively. The numbers of the regions included in each sub - map in the second figure and the third figure remain unchanged, and the size of the overlapping region 5 is still ; After the third - layer convolution operation, the fourth figure is obtained, that is, the output feature map; among them, the first sub - map includes region 1, the second sub - map includes region 2, the third sub - map includes region 3, the fourth sub - map includes region 4, and there is no overlapping region between adjacent sub - maps.

[0062] Specifically, Figure 3 and Figure 4 , each small square in represents a data. Therefore, by counting the number of small squares in each sub - map, the sum of the data amounts required for convolution operation of each sub - map can be obtained, and the computational amount of the convolution operation is proportional to the data amount. Thus, the computational amount of the convolution operation in the related technology and the computational amount of the convolution operation in the embodiment of the present application can be intuitively compared.

[0063] As Figure 3 shown, the sum of the data amounts of the first sub - map to the fourth sub - map in the first - layer convolution operation is: ; the sum of the data amounts of the first sub - map to the fourth sub - map in the second - layer convolution operation is: ; the sum of the data amounts of the first sub - map to the fourth sub - map in the third - layer convolution operation is: ; It can be seen from this that Figure 3The sum of the data volumes of the first to fourth sub - graphs in the middle three - layer convolution operation is: 324 + 196 + 100 = 620. As Figure 4 shown, the sum of the data volumes of the first to fourth sub - graphs in the first - layer convolution operation is: ; the sum of the data volumes of the first to fourth sub - graphs in the second - layer convolution operation is: ; the sum of the data volumes of the first to fourth sub - graphs in the third - layer convolution operation is: ; thus, it can be known that Figure 4 the total data volume of each sub - graph in the middle three - layer convolution operation is: 196 + 144 + 100 = 440. By comparison, when splitting the input feature map into sub - graphs, only according to the convolution kernel size and stride of the current convolution layer to determine the size of the overlapping area can effectively reduce the size of the overlapping area, and thus effectively reduce the data volume.

[0064] Furthermore, Figure 3 and Figure 4 when calculating the sum of the data volumes of each sub - graph in each layer of convolution operation, the data volumes of each sub - graph are superimposed to obtain the sum of the data volumes of the first to fourth sub - graphs; for the embodiments of the present application, the overlapping areas of each sub - graph in each layer of convolution operation are not calculated repeatedly. In other words, if the sub - graph of the prior convolution operation already includes the area overlapping with the sub - graph of the subsequent convolution operation, the sub - graph of the subsequent convolution operation does not calculate this overlapping area repeatedly. Therefore, Figure 4 in, the sum of the data volumes of the first to fourth sub - graphs in each layer of convolution operation is the number of small squares of the corresponding feature map; in the first - layer convolution operation, the input feature map is a square, so the data volume of the first to fourth sub - graphs is ; similarly, it can be known that in the second - layer convolution operation, the data volume of the first to fourth sub - graphs is ; in the third - layer convolution operation, the data volume of the first to fourth sub - graphs is ; therefore, Figure 4 the total data volume of each sub - graph in the middle three - layer convolution operation is: 144 + 100 + 64 = 308; that is to say, when only considering the convolution kernel size and stride of the current convolution layer to split the input feature map, the sub - graph of the subsequent convolution operation does not calculate the overlapping area in the sub - graph of the prior convolution operation repeatedly, which can further reduce the total data volume.

[0065] Based on the above - mentioned embodiments, as an alternative embodiment, before the sub - graph stores the operation result of the area to be calculated and the operation result of the second area respectively, it further includes: Determine the coordinates of the second area in the sub - graph, and obtain the operation result of the second area from the operation result of the area to be calculated in the sub - graph according to the coordinates.

[0066] In the embodiments of the present application, according to the position of the second region in the original graph and the sub - graph division method, the coordinates of the second region in the corresponding sub - graph can be calculated; after obtaining the coordinates of the second region in the sub - graph, the operation result of the second region is extracted from the operation results of the regions to be calculated in the sub - graph.

[0067] Optionally, in the embodiments of the present application, when there is an overlapping region between the sub - graph of the prior convolution operation and multiple sub - graphs of the subsequent convolution operations, it is necessary to determine the starting coordinates of the overlapping region between the sub - graph of the prior convolution operation and each sub - graph of the subsequent convolution operations; that is to say, when the second region of the sub - graph of the prior convolution operation includes the regions overlapping multiple sub - graphs of the subsequent convolution operations, it is necessary to determine the starting coordinates of the overlapping region of each sub - graph of the subsequent convolution operations in the sub - graph of the prior convolution operation, so as to facilitate obtaining the operation results of the overlapping regions of each sub - graph of the subsequent convolution operations from the operation results of the regions to be calculated in the sub - graph of the prior convolution operation.

[0068] In the embodiments of the present application, by determining the coordinates of the second region in the sub - graph and extracting the operation result of the second region according to the coordinates, the operation result of the second region can be saved separately, and the sub - graph of the subsequent convolution operation can obtain the operation results of the overlapping regions with the sub - graph of the prior convolution operation according to the operation result of the second region of the sub - graph of the prior convolution operation, avoiding repeated calculations of the overlapping regions between sub - graphs. This precise data access method helps to reduce calculation redundancy and improve calculation efficiency.

[0069] Based on the above embodiments, as an optional embodiment, the sub - graph stores the operation results of the regions to be calculated and the operation results of the second region respectively, including: For the sub - graph with a non - empty second region, the operation result of the region to be calculated includes the operation result of the first region and the operation result of the second region; the operation result of the region to be calculated and the operation result of the second region are stored in their respective corresponding storage spaces; For the sub - graph with an empty second region, the operation result of the region to be calculated includes the operation result of the first region, and the operation result of the first region is stored in the corresponding storage space.

[0070] In the embodiments of the present application, for each sub - graph, first, it is judged whether its second region is empty. If the second region is not empty, then the operation result of the region to be calculated will include the operation result of the first region and the operation result of the second region, and these two parts of operation results are stored in their respective corresponding storage spaces; if the second region is empty, then the operation result of the region to be calculated only includes the operation result of the first region, and this part of the operation result is stored in the corresponding storage space. Since the second region does not exist, no additional storage space needs to be allocated for it.

[0071] Optionally, in the embodiments of the present application, if the operation result of the second region includes the operation results of multiple sub-graph overlapping regions in the subsequent convolution operations, the operation result of the second region can be split into the operation results of each sub-graph overlapping region in the subsequent convolution operations, and the split operation results can be stored separately.

[0072] In the embodiments of the present application, by storing the operation result of the region to be calculated and the operation result of the second region separately, it is convenient for the sub-graphs in the subsequent convolution operations to obtain the operation results of the overlapping regions they need from the storage space, avoiding repeated calculations of the overlapping regions by the sub-graphs in the subsequent convolution operations and saving computing resources; at the same time, for the sub-graphs with an empty second region, only the operation result of the first region is stored, avoiding unnecessary waste of storage space.

[0073] Based on the above embodiments, as an alternative embodiment, when the number of convolution operation layers in the region to be calculated of the sub-graph is one layer; Obtaining the operation result of the sub-graph from the storage space to generate an output feature map includes: Reading the operation result of the region to be calculated of the sub-graph from the storage space, and splicing the operation results of the region to be calculated of the sub-graph to generate an output feature map.

[0074] In the embodiments of the present application, when the number of convolution operation layers is one layer, the results after one layer of convolution operation of the region to be calculated of the sub-graph are read from the storage space (such as memory or hard disk), and these results are usually stored in the form of matrices or tensors, representing the feature responses of the sub-graph at different positions; according to the positions and mutual relationships of the sub-graphs in the original image, the operation results of each sub-graph are spliced. The splicing method is usually based on the spatial position, that is, combined according to the arrangement order and relative position of the sub-graphs in the original image, and the spliced result is used as the output feature map.

[0075] In the embodiments of the present application, when the number of convolution operation layers is one layer, by splicing the operation results of the sub-graphs, the output feature map can integrate information from different sub-graphs, which helps the model to more comprehensively understand the content of the input image and improve the accuracy and robustness of recognition.

[0076] Based on the above embodiments, as an alternative embodiment, when the number of convolution operation layers in the region to be calculated of the sub-graph is multiple layers; Obtaining the operation result of the sub-graph from the storage space to generate an output feature map includes: For non - the last - layer convolution operation, the sub - graph of the prior convolution operation obtains the first operation result from the storage space, and uses the first operation result as the input data for the next convolution layer of the sub - graph of the prior convolution operation; wherein, the first operation result is the operation result of the sub - graph of the prior convolution operation in the area to be calculated in the current convolution layer; the sub - graph of the subsequent convolution operation obtains the second operation result and the third operation result from the storage space; wherein, the second operation result is the operation result of the sub - graph of the subsequent convolution operation in the area to be calculated in the current convolution layer, and the third operation result is the operation result of the sub - graph of the prior convolution operation in the second area in the current convolution layer. The second operation result and the third operation result are concatenated as the input data for the next convolution layer of the sub - graph of the subsequent convolution operation. For the last - layer convolution operation, the operation result of the area to be calculated of the sub - graph of the prior convolution operation and the operation result to be calculated of the sub - graph of the subsequent convolution operation are concatenated to generate the output feature map.

[0077] Specifically, when there are multiple layers of convolution operations, since the area to be calculated of the sub - graph of the subsequent convolution operation does not include the area overlapping with the sub - graph of the prior convolution operation, the operation result of the area to be calculated of the sub - graph of the subsequent convolution operation does not include the operation results of all areas of this sub - graph. Therefore, it is necessary to obtain the operation results of the overlapping area of this sub - graph from the storage space, and concatenate the operation result of the area to be calculated of the sub - graph of the subsequent convolution operation and the operation results of the overlapping area as the input data for the next - layer convolution operation of the sub - graph of the subsequent convolution operation.

[0078] Specifically, for the last - layer convolution operation, since there is no next - layer convolution operation, it is not necessary to determine the input data for the next convolution layer; in the embodiments of the present application, the operation results of the areas to be calculated of the sub - graphs are obtained from the storage space, and the operation results of the areas to be calculated of the sub - graphs are concatenated to generate the output feature map.

[0079] In the embodiments of the present application, the sub - graph of the prior convolution operation stores the operation results of the overlapping area in the storage space, so that the sub - graph of the subsequent convolution operation can obtain the operation results of the overlapping area from the storage space, and concatenate the operation result of the area to be calculated of the sub - graph of the subsequent convolution operation and the operation results of the overlapping area as the input data for the next - layer convolution operation of the sub - graph of the subsequent convolution operation, thereby ensuring that the subsequent sub - graph can obtain the input data for the next - layer convolution operation without calculating the overlapping area, effectively saving computing resources and reducing computing time.

[0080] Figure 5 As shown in the figure, it is a schematic structural diagram of a data processing device provided by the embodiments of the present application. The device may include: a first processing module 501, a second processing module 502, a third processing module 503, and a fourth processing module 504.

[0081] The first processing module 501 is configured to obtain a plurality of sub - graphs obtained by splitting an input feature map; wherein, there is an overlapping area between adjacent sub - graphs; The second processing module 502 is configured to determine the area to be calculated of the sub - graphs according to the order of convolutional operations of the sub - graphs; wherein, the area to be calculated includes a first area and a second area; the first area includes the part that does not overlap with other sub - graphs, and the second area includes the part that overlaps with other sub - graphs; and the second area of the sub - graph that undergoes the prior convolutional operation includes: the overlapping area that overlaps with the sub - graph that undergoes the subsequent convolutional operation; and the second area of the sub - graph that undergoes the subsequent convolutional operation does not include: the area that overlaps with the sub - graph that undergoes the prior convolutional operation; The third processing module 503 is configured to sequentially perform convolutional operations on the area to be calculated of the sub - graphs and cache the calculation results in a storage space; wherein, the calculation results of the first area and the second area are stored separately; The fourth processing module 504 is configured to obtain the output data of the sub - graphs from the storage space and generate an output feature map.

[0082] The data processing device according to the embodiment of the present application can execute the data processing method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the data processing device according to the embodiments of the present application correspond to the steps in the data processing method of the embodiments of the present application. For the detailed function descriptions of each module of the data processing device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.

[0083] In the embodiment of the present application, by dividing the area of the sub - graphs into overlapping areas and non - overlapping areas and according to the order of convolutional operations of the sub - graphs, the area to be calculated of the sub - graph that undergoes the subsequent convolutional operation does not include the area that overlaps with the sub - graph that undergoes the prior convolutional operation. In this way, the sub - graph that undergoes the subsequent convolutional operation does not need to calculate the area that overlaps with the sub - graph that undergoes the prior convolutional operation; at the same time, the sub - graph that undergoes the prior convolutional operation stores the calculation results of the overlapping area in the storage space, so that the sub - graph that undergoes the subsequent convolutional operation can directly obtain the calculation results of the overlapping area from the storage space, ensuring that the overlapping area between sub - graphs will not be repeatedly calculated, saving computing resources and reducing computing time.

[0084] Further, in the embodiment of the present application, the parameters of the convolutional operation can also be preset in a software - configured manner, including but not limited to: the number of convolutional layers, the size of the convolutional kernel for each convolutional operation, the stride, and the number of sub - graphs; according to the software - configured parameters, the size of the overlapping area of each convolutional layer is determined without considering the size of the overlapping area in the next - layer convolutional operation, which can effectively reduce the size of the overlapping area and further effectively reduce the overlapping data.

[0085] Figure 6 It is a schematic structural diagram of an electronic device provided by the embodiment of the present application, asFigure 6 As shown in Figure 6 , the electronic device 4000 includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0086] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present application. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0087] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0088] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0089] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0090] Among them, the electronic device package may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0091] The embodiments of the present application provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented. Compared with the prior art, it can achieve: It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0092] The embodiments of the present application also provide a computer program product, including a computer program, and when the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented. Compared with the prior art, it can achieve: The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims, and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0093] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0094] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A data processing method, applied to a neural network processing unit NPU, characterized in that: The method comprises: Obtain multiple sub-images divided from the input feature map; wherein there are overlapping areas between adjacent sub-images; According to the order of the sub-graph convolution operation, the area to be calculated of the sub-graph is determined; wherein the area to be calculated includes a first area and a second area; the first area includes a portion that does not overlap with other sub-graphs, and the second area includes a portion that overlaps with other sub-graphs; and the second area of ​​the sub-graph of the previous convolution operation includes: an overlapping area that overlaps with the sub-graph of the subsequent convolution operation; and the second area of ​​the sub-graph of the subsequent convolution operation does not include: an area that overlaps with the sub-graph of the previous convolution operation; Convolution operations are sequentially performed on the to-be-calculated area of ​​the sub-graph, and the operation results are cached in a storage space; wherein the operation results of the to-be-calculated area of ​​the sub-graph and the operation results of the second area are stored separately; The operation result of the sub-graph is obtained from the storage space to generate an output feature graph.

2. The data processing method according to claim 1, characterized in that: The convolution operation of the input feature map uses a preset convolution kernel and step size; The step of obtaining the plurality of sub-graphs obtained by segmenting the input feature graph further includes: According to the preset convolution kernel size and the step size, the input feature map is divided into a plurality of sub-maps.

3. The data processing method according to claim 2, characterized in that: The step of dividing the input feature map into a plurality of sub-maps according to the preset convolution kernel size and the step size includes: According to the preset convolution kernel and convolution step size, determine the size of the overlapping area when the convolution kernel in the corresponding convolution layer slides as the first size; The size of the overlapping area of ​​each sub-image is determined according to the first size, and the input feature image is divided into multiple sub-images according to the size of the overlapping area of ​​each sub-image.

4. The data processing method according to claim 3, characterized in that: The subgraph stores the calculation result of the area to be calculated and the calculation result of the second area respectively, and also includes: The coordinates of the second area in the sub-graph are determined, and the calculation result of the second area is obtained from the calculation result of the area to be calculated of the sub-graph according to the coordinates.

5. The data processing method according to claim 4, characterized in that: The subgraph stores the calculation result of the area to be calculated and the calculation result of the second area respectively, including: For a subgraph in which the second region is not empty, the calculation result of the region to be calculated includes the calculation result of the first region and the calculation result of the second region; the calculation result of the region to be calculated and the calculation result of the second region are stored in their respective corresponding storage spaces; For a subgraph in which the second region is empty, the calculation result of the region to be calculated includes the calculation result of the first region, and the calculation result of the first region is stored in a corresponding storage space.

6. The data processing method according to claim 5, characterized in that: When the number of layers of the convolution operation of the to-be-calculated area of ​​the subgraph is one layer; The obtaining the operation result of the sub-graph from the storage space to generate an output feature graph includes: The calculation results of the to-be-calculated area of ​​the sub-graph are read from the storage space, and the calculation results of the to-be-calculated area of ​​the sub-graph are concatenated to generate an output feature map.

7. The data processing method according to claim 5, characterized in that: When the number of layers of the convolution operation of the to-be-calculated area of ​​the subgraph is multiple layers; The obtaining the operation result of the sub-graph from the storage space to generate an output feature graph includes: For a non-last convolution operation, the subgraph of the previous convolution operation obtains a first operation result from the storage space, and uses the first operation result as input data for the next convolution layer of the subgraph of the previous convolution operation; wherein the first operation result is the operation result of the area to be calculated of the subgraph of the previous convolution operation in the current convolution layer; the subgraph of the subsequent convolution operation obtains a second operation result and a third operation result from the storage space; wherein the second operation result is the operation result of the area to be calculated of the subgraph of the subsequent convolution operation in the current convolution layer, and the third operation result is the operation result of the second area of ​​the subgraph of the previous convolution operation in the current convolution layer, and the second operation result and the third operation result are concatenated as input data for the next convolution layer of the subgraph of the subsequent convolution operation; For the last layer of convolution operation, the operation results of the to-be-calculated area of ​​the sub-graph of the previous convolution operation and the to-be-calculated results of the sub-graph of the subsequent convolution operation are concatenated to generate an output feature map.

8. A data processing device, characterized in that: The device is applied to a neural network processing unit NPU, and the device includes: The first processing module is used to obtain a plurality of sub-images segmented from the input feature image, wherein there are overlapping areas between adjacent sub-images; A second processing module is used to determine the area to be calculated of the sub-graph according to the order of the sub-graph convolution operation; wherein the area to be calculated includes a first area and a second area; the first area includes a portion that does not overlap with other sub-graphs, and the second area includes a portion that overlaps with other sub-graphs; and the second area of ​​the sub-graph in the previous convolution operation includes: an overlapping area with the sub-graph in the subsequent convolution operation; and the second area of ​​the sub-graph in the subsequent convolution operation does not include: an area overlapping with the sub-graph in the previous convolution operation; A third processing module is used to perform convolution operations on the to-be-calculated areas of the subgraph in sequence, and cache the calculation results in a storage space; wherein the calculation results of the first area and the calculation results of the second area are stored separately; The fourth processing module is used to obtain the output data of the sub-graph from the storage space to generate the output feature graph.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.