Visual Model Inference Method and Device Based on Image Segmentation

By dividing the image and video memory into slices and blocks, and using the block list index to stitch the image slices, the problem of low memory utilization of the visual model when processing dynamic-sized images is solved, and resource utilization and inference efficiency are improved.

CN120144065BActive Publication Date: 2025-07-22LONGSHINE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510623786.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the prior art, the visual model has low memory utilization rate when processing dynamic-sized images, resulting in waste of resources, which is particularly prominent in mobile devices.

Method used

The images to be processed and physical video memory are divided into multiple image slices and video memory blocks, indexes are recorded through block lists, and image slices are stitched by index when visual model reasoning, realizing fixed-size storage and utilization.

Benefits of technology

It improves the utilization rate of video memory, avoids the waste of pre-allocated video memory, and improves the inference efficiency of the visual model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144065B_ABST
    Figure CN120144065B_ABST
Patent Text Reader

Abstract

The present invention provides a visual model inference method and apparatus based on image segmentation, relating to the technical field of image processing. The method includes: segmenting an image to be processed according to a preset size to obtain a plurality of image slices, recording the corresponding image indexes in a block list, and dividing the physical video memory storing the image to be processed into a plurality of video memory blocks to store the image slices. When the visual model performs inference, the image slices are obtained from the video memory blocks according to the block list and spliced into an image to be inferred, so as to obtain an image inference result. The present invention divides the image to be processed and the physical video memory respectively, stores the segmented image to be processed through the divided video memory blocks, and records them in the block list for the visual model to perform inference, solving the problem of low video memory utilization rate when the visual model in the prior art performs inference on images of various dynamic sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a visual model inference method and device based on image segmentation. Background Art

[0002] In the field of computer vision, deep learning models have been widely applied to tasks such as image recognition, classification, and segmentation. When a visual model is trained, the input images generally have a fixed size, and the batch size and the size of intermediate calculation results (such as feature maps) are also fixed. However, during inference, the input batch size of the model is not fixed, and the size of each picture is not necessarily fixed either. Moreover, during the long-term inference of the entire model, the input size may also change dynamically. If the model can only receive inputs of a fixed size, then it cannot be used, and the inference engine needs to be adjusted and optimized accordingly to adapt to the inference with dynamic shapes. When the shape of the input changes, the video memory allocation during the model calculation process is a difficult problem. If the maximum size of the picture is directly preset, video memory fragmentation will occur, resulting in a problem of wasted video memory, and this problem will be more prominent in mobile devices with a smaller memory capacity.

[0003] To solve the above technical problems, in the prior art, in order to support dynamic shapes, a deep learning inference engine of TensorRT is proposed. It introduces an "optimized configuration file", which enables users to specify a series of possible input tensor size ranges according to this optimized configuration file. However, TensorRT still pre-allocates video memory according to the maximum size in the optimized configuration file, resulting in a still low video memory utilization rate. If the actual input size is much smaller than the specified maximum size, then most of the pre-allocated video memory will not be fully utilized, causing resource waste. Summary of the Invention

[0004] The present invention provides a visual model inference method and device based on image segmentation, which are used to solve the problem of low video memory utilization rate when the visual model in the prior art infers images of various dynamic sizes.

[0005] The present invention provides a visual model inference method based on image segmentation, which is characterized by including:

[0006] Segment the image to be processed according to a preset size to obtain a plurality of image slices, and record the image index of each image slice in a block list;

[0007] Divide the physical video memory storing the image to be processed into a plurality of video memory blocks, and store the image slices through the video memory blocks;

[0008] When the visual model performs inference, the image slices are obtained from the video memory blocks according to the block list, and the image slices are stitched together according to the image index to obtain the image to be inferred;

[0009] The image to be inferred is input into the visual model for inference to obtain an image inference result.

[0010] In some embodiments, storing the image slices in the video memory blocks includes:

[0011] Combining the image slices into image blocks according to a preset quantity;

[0012] Recording the image blocks and the number of image slices included in the image blocks in the block list, and establishing a mapping relationship between the image blocks and the video memory blocks in the block list;

[0013] Storing the image slices in the image blocks into the video memory blocks according to the mapping relationship.

[0014] In some embodiments, obtaining the image slices from the video memory blocks according to the block list includes:

[0015] Querying the mapping relationship between the image blocks and the video memory blocks from the block list according to the image index of the image slices;

[0016] Determining the target video memory block corresponding to the storage of the image slices from the physical video memory according to the mapping relationship, and obtaining the image slices with the corresponding image index according to the target video memory block.

[0017] In some embodiments, inputting the image to be inferred into the visual model for inference to obtain an image inference result includes:

[0018] During the process of the visual model extracting the feature map of the image to be inferred, obtaining the feature maps extracted by each layer of the network, and splitting the feature maps to obtain a plurality of feature sub - maps;

[0019] Recording the feature indexes of the feature sub - maps in the block list, and storing the feature sub - maps in the video memory blocks divided by the physical video memory;

[0020] When the visual model performs inference, querying the feature sub - maps from the block list according to the feature indexes;

[0021] Performing a stitching process on the feature sub - maps according to the feature indexes to obtain a feature map to be processed;

[0022] Performing a mapping process on the feature map to be processed output by the last layer of the network in the visual model to obtain an image inference result.

[0023] In some embodiments, before segmenting the image to be processed according to a preset size, the method further includes:

[0024] Performing size filling processing on the image to be processed, and constructing a filling configuration file according to the size filling result;

[0025] The filling configuration file is used to perform size cropping processing on the predicted feature map according to the size filling result before the visual model outputs the image inference result, and the predicted feature map is the feature map used for mapping processing to obtain the image inference result.

[0026] In some embodiments, after inputting the image to be inferred into the visual model for inference to obtain an image inference result, the method further includes:

[0027] Removing the image index corresponding to the image slice of the image to be inferred and the feature index corresponding to the feature submap from the block list;

[0028] Releasing the storage space of the video memory block used to store the image slice and the feature submap in the physical video memory.

[0029] A visual model inference device based on image segmentation according to the present invention includes:

[0030] A segmentation module, configured to segment the image to be processed according to a preset size to obtain a plurality of image slices, and record the image index of each image slice in a block list;

[0031] A storage module, configured to divide the physical video memory storing the image to be processed into a plurality of video memory blocks, and store the image slices through the video memory blocks;

[0032] An acquisition module, configured to, when the visual model performs inference, acquire the image slices from the video memory blocks according to the block list, and splice the image slices according to the image index to obtain an image to be inferred;

[0033] An inference module, configured to input the image to be inferred into the visual model for inference to obtain an image inference result.

[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method for visual model inference based on image segmentation as described in any one of the above is implemented.

[0035] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for visual model inference based on image segmentation as described in any one of the above is implemented.

[0036] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for visual model inference based on image segmentation as described in any one of the above is implemented.

[0037] The method and device for visual model inference based on image segmentation provided by the present invention divide a to-be-processed image and physical video memory respectively, and then use the divided video memory blocks to store the segmented to-be-processed image, which is recorded in a block list for the visual model to perform inference. In this way, even for to-be-processed images with different dynamic sizes, they can be segmented according to a fixed size and stored in the divided video memory blocks as needed, effectively utilizing the video memory resources without pre-allocation in advance, and improving the utilization rate of the video memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 is a schematic flowchart of the method for visual model inference based on image segmentation provided by the present invention.

[0040] Figure 2 is a schematic diagram of the segmentation of the to-be-processed image provided by the present invention.

[0041] Figure 3 is a schematic diagram of the block list provided by the present invention.

[0042] Figure 4 is a schematic diagram of the visual model inference process provided by the present invention.

[0043] Figure 5 is a schematic structural diagram of the device for visual model inference based on image segmentation provided by the present invention.

[0044] Figure 6 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] The visual model inference method and apparatus based on image segmentation of the present invention will be described below with reference to the accompanying drawings. Figure 1 is a schematic flowchart of the visual model inference method based on image segmentation provided by the present invention, as Figure 1 shown. The method includes the following steps 101 to 104, which will be specifically described below.

[0047] Step 101: Segment the image to be processed according to a preset size to obtain a plurality of image slices, and record the image index of each image slice in a block list.

[0048] Here, there are multiple images to be processed that need to be inferred using a visual model, and the image sizes of each image to be processed are different. The image size generally refers to the width and height of the image. Therefore, for any image to be processed with any size, the image to be processed is segmented according to the preset size to obtain a plurality of image slices. The preset size can be, for example, 32×32. When the image to be processed is smaller than the preset size, it can be filled with blank pixels to the preset size or an integer multiple of the preset size.

[0049] Next, the image index of each image slice is recorded in the block list. This image index is set according to the order of the image slices and is used to query the corresponding image slice, and can also be used to represent the position in the image. Here, in the embodiment of the present invention, a block list is constructed to record the corresponding image index for subsequent storage of the image slices.

[0050] For example Figure 2 shown, an image to be processed is segmented according to the preset size, and finally 15 image slices are obtained, and image indices 1, 2, 3,..., 15 are set for each image slice.

[0051] Step 102: Divide the physical video memory storing the image to be processed into multiple video memory blocks, and store the image slices through the video memory blocks.

[0052] Before performing inference on the image to be processed, it is generally necessary to allocate corresponding physical video memory for storage, such as the memory of the computer's graphics card device or graphics processing unit. Similar to the segmentation process of the image to be processed, the physical video memory storing the image to be processed is also divided into multiple video memory blocks, denoted as block. Of course, the memory allocation size of the video memory block is also preset, and the memory allocation size of the video memory block can be reasonably preset according to the model structure of the vision model, the input characteristics of the model, the model calculation precision type, etc. Then, the segmented image slices are stored one by one through the divided video memory blocks.

[0053] For example Figure 3 As shown, the physical video memory storing the image to be processed is divided into 7 video memory blocks, such as video memory block 1, video memory block 2, …, video memory block 7, and 15 image slices obtained by segmenting the image to be processed are respectively stored in video memory blocks 1, 2, 4, and 6 among them, while the video memory blocks 3, 5, and 7 in the physical video memory where no image slices are stored are in an idle state.

[0054] Step 103: When the vision model performs inference, obtain the image slices from the video memory blocks according to the block list, and splice the image slices according to the image index to obtain the image to be inferred.

[0055] When the vision model performs inference, it is necessary to obtain the image to be processed from the physical video memory. Therefore, the image slices are obtained from the video memory blocks according to the block list here. The block list can be queried by using the image index of the image slice, and then the video memory blocks of the physical video memory are traversed to determine the video memory block where the image slice is stored, so as to obtain the corresponding image slice. Then, the image slices can be spliced according to the image index to obtain the image to be inferred. The image index can be used to represent the position of the image slice in the image. Therefore, splicing according to the image index can restore a complete image as the image to be inferred.

[0056] Step 104: Input the image to be inferred into the vision model for inference to obtain the image inference result.

[0057] Finally, perform image inference. Input the image to be inferred into the vision model for inference to obtain the image inference result. The vision model can be a convolutional neural network, a Transformer model, or other models or encoders used for image processing. The inference process can be image segmentation, object detection, or image reconstruction, etc. The finally obtained image inference result can be the corresponding image classification category, image segmentation result, or the object detected in the image, etc.

[0058] In the embodiments of the present invention, by dividing the image to be processed and the physical video memory respectively, and then using the divided video memory blocks to store the sliced images to be processed, which are recorded in the block list for the visual model to perform inference. In this way, even for images to be processed with different dynamic sizes, they can be sliced according to a fixed size and stored in the divided video memory blocks as needed, effectively utilizing the video memory resources without pre-allocation in advance, and improving the utilization rate of the video memory.

[0059] In some embodiments, the storage process of storing image slices in video memory blocks can be implemented through the following process, which specifically includes: first combining image slices into image blocks according to a preset quantity, recording the image blocks and the number of image slices included in the image blocks in the block list, and establishing a mapping relationship between the image blocks and the video memory blocks in the block list. Finally, storing the image slices in the image blocks into the video memory blocks according to the mapping relationship.

[0060] Here, since the number of image slices may be very large, it is impossible to allocate a video memory block for each image slice to store. Therefore, here the image slices are first combined into image blocks according to a preset quantity, and in this way, video memory blocks are allocated for storage with the image blocks as the storage unit.

[0061] As Figure 2 shown, the image to be processed is sliced to obtain 15 image slices, and then combined into 4 image blocks according to 4 slices per block. Image slices 1 to 4, 5 to 8, and 9 to 12 are respectively combined into image blocks 1, 2, and 3, and the remaining 3 image slices 13 to 15 are combined into image block 4.

[0062] Then record the image blocks and the number of image slices included in the image blocks in the block list. As Figure 3 shown, when storing images, image blocks 1, 2, 3, and 4 are respectively recorded in the block list, and the number of image slices corresponding to each image block is also recorded in the filled column of the block list. Here, it is also necessary to establish a mapping relationship between the image blocks and the video memory blocks in the block list. Since video memory blocks are allocated for storage with the image blocks as the storage unit, a mapping relationship needs to be established for each image block. The video memory block is the actual address for storing image slices, and the actual address of the image slices can be conveniently queried through the mapping relationship during query. For example Figure 3 shown, the mapping relationship between the image blocks and the video memory blocks is recorded in the physical block number column of the block list. The image blocks 1, 2, 3, and 4 combined from the image slices of the image to be processed respectively map to video memory blocks 1, 4, 6, and 2 of the physical video memory.

[0063] Finally, store the image slices in the image block into the video memory block according to the mapping relationship. After the mapping relationship is determined, image storage can be performed to store the image slices in the image block into the corresponding video memory block. As Figure 3 shown, store image blocks 1, 2, 3, and 4 into video memory blocks 1, 4, 6, and 2 of the physical video memory respectively.

[0064] In the embodiment of the present invention, the image slices are combined into image blocks according to a reasonable quantity, and the video memory blocks are allocated in units of image blocks. A mapping relationship between the image blocks and the video memory blocks is established for image storage. In this way, there is no need to allocate video memory blocks for each image slice, which can effectively save video memory resources. By establishing a mapping relationship between the image and the physical video memory, fast query and acquisition of image slices can be achieved during model inference.

[0065] In some embodiments, the process of obtaining image slices from the video memory block according to the block list can be implemented through the following process. First, query the mapping relationship between the image block and the video memory block from the block list according to the image index of the image slice. Then, according to the mapping relationship, determine the target video memory block where the image slice is stored in the physical video memory, and obtain the image slice corresponding to the image index according to the target video memory block.

[0066] When the visual model performs inference, at this time, the corresponding image to be inferred needs to be obtained from the physical video memory. This requires querying the corresponding image slice in the physical video memory. First, query the block list according to the image index. According to the image index, the image block corresponding to the image slice can be determined. Then, query the physical block number column of the block list to find the mapping relationship between the image block and the video memory block, so as to determine the corresponding target video memory block in the physical video memory. By traversing the target video memory block, the image slice corresponding to the image index can be queried.

[0067] For example, as Figure 3 shown, for the image slice with the image index of 7, first determine that the image block where it is located is image block 2. Then query the physical block number column to determine that there is a mapping relationship between image block 2 and video memory block 4 in the physical video memory. Video memory block 4 is the target video memory block. Finally, in video memory block 4, the corresponding image slice can be obtained according to the image index 7. For each image slice, a similar method can be used to obtain the corresponding image slice from the physical video memory one by one.

[0068] In the embodiment of the present invention, when obtaining image slices from the physical video memory, by using the preset image index and the mapping relationship established in the block list, the image slices can be quickly queried and obtained, which are used to splice into the image to be inferred for inference, and the acquisition efficiency of the image can be improved.

[0069] In some embodiments, when the image to be inferred is input into the vision model for inference, some intermediate activation values during the inference process can also be stored in the divided video memory blocks. The intermediate activation values can be, for example, the feature maps of the image to be inferred, that is, the feature maps output by each layer of the network in the vision model. The following specifically introduces the specific process when the vision model infers the image to be inferred.

[0070] First, during the process of the vision model extracting the feature maps of the image to be inferred, the feature maps extracted by each layer of the network are obtained, and the feature maps are sliced to obtain multiple feature sub-maps.

[0071] As Figure 4 shown, the vision model is, for example, obtained by connecting multiple convolutional layer networks in series. After the image to be inferred is input into the vision model, each convolutional layer network will output the corresponding feature map, and these feature maps also need to be allocated corresponding physical video memory for storage.

[0072] Next, similar to the image to be processed, the feature maps are sliced to obtain multiple feature sub-maps. Since the sizes of the feature maps obtained by each layer of the network are different, they cannot be sliced according to a unified size. Therefore, here, according to the size of the feature map obtained by each layer of the network, the slicing size is reasonably preset to ensure that the number of feature sub-maps obtained by slicing each feature map is the same as the number of image slices obtained when the image to be processed is sliced. For example, if the image to be processed is sliced into 9 image slices, then the feature maps obtained by each layer of the network are also sliced into 9 feature sub-maps.

[0073] Similarly, here, a corresponding feature index is also set for each sliced feature sub-map, the feature index of the feature sub-map is recorded in the block list, and the feature sub-map is stored through the video memory blocks divided by the physical video memory. The storage method of the feature sub-map is similar to that of the image slice. First, the feature sub-maps are combined into feature blocks according to the preset number, then the feature blocks and the number of feature sub-maps included in the feature blocks are recorded in the block list, and a mapping relationship between the feature blocks and the video memory blocks is established in the block list. Finally, the feature sub-maps in the feature blocks are stored in the video memory blocks according to this mapping relationship.

[0074] When the vision model performs inference, the feature sub-maps are queried from the block list according to the feature index. Similarly, the query and acquisition of the feature sub-maps are similar to those of the image slices. First, according to the feature index of the feature sub-map, the mapping relationship between the feature blocks and the video memory blocks is queried from the block list, then according to the mapping relationship, the target video memory block corresponding to the storage of the feature sub-map is determined from the physical video memory, and the feature sub-map with the corresponding feature index is obtained according to the target video memory block.

[0075] After obtaining the feature sub - graphs, the feature sub - graphs are then stitched according to the feature index to obtain the feature graph to be processed. The stitching process of the feature sub - graphs is similar to image slicing, which will not be elaborated here.

[0076] Finally, the feature graph to be processed output by the last layer of the vision model is mapped to obtain the image inference result.

[0077] The feature graphs output by each layer of the vision model are stored in the physical video memory in a similar manner as above. During the processing of the next layer of the network, the corresponding feature graph to be processed is obtained from the physical video memory for processing, and the corresponding feature graph to be processed output after processing is also input into the physical video memory. Until the last layer of the vision model, the feature graph to be processed output by the last layer of the vision model is mapped to obtain the image inference result. The mapping process can be achieved through the activation function of a fully - connected layer neural network or a multi - layer perceptron. The image inference result can be the category of the corresponding image classification, the image segmentation result, or the objects detected in the image, etc.

[0078] As Figure 4 shown, a certain image to be processed is sliced into 9 image slices, namely image slices 1 to 9, and then stored in the corresponding video memory blocks of the physical video memory respectively. When the vision model performs inference, it queries from the physical video memory to obtain 9 image slices, and inputs the stitched image to be inferred into the first convolutional layer network of the vision model, that is, convolutional layer network 1. Then the feature graph output by convolutional layer network 1 is also sliced into 9 feature sub - graphs, which are stored in the corresponding video memory blocks of the physical video memory in a similar manner as the image slices. After that, continue to query from the physical video memory in a similar manner to obtain 9 feature sub - graphs, stitch them to obtain the corresponding feature graph to be processed, and input it into the second convolutional layer network of the vision model for processing, that is, convolutional layer network 2. And so on, until the corresponding feature graph to be processed is output by the Nth convolutional layer network (i.e., convolutional layer network N). Here, the feature graph to be processed output by the last convolutional layer network of the vision model, that is, the feature graph to be processed output by convolutional layer network N, is mapped to obtain the corresponding image inference result.

[0079] In the embodiment of the present invention, in a similar manner to the image slices obtained by slicing the image to be processed, the feature graphs output by each convolutional layer network of the vision model are also sliced and stored in the divided video memory blocks. In this way, during the inference process of the vision model, the utilization rate of the video memory can be effectively improved, and the inference efficiency of the vision model can be improved.

[0080] In some embodiments, considering that the images to be processed have various different sizes and it is difficult to preset the segmentation size, before segmenting the images to be processed according to the preset size, the embodiments of the present invention also uniformly perform padding processing on the images to be processed to ensure that they can be segmented according to the preset size.

[0081] Specifically, perform size padding processing on the images to be processed and construct a padding configuration file according to the size padding result. For each image to be processed, size padding processing is uniformly performed. For example, the preset slice size of the images to be processed is 32×32, and the sizes of some images to be processed are less than 32×32 or not an integer multiple of 32×32. In this case, the images to be processed cannot be reasonably segmented according to 32×32. Therefore, through size padding processing here, it is ensured that the size of each image to be processed is an integer multiple of 32×32. The padding processing method can be to fill with blank pixels to increase the width and height of the images to be processed, so that the size increases to an integer multiple of 32×32.

[0082] Since the blank pixels filled in the images to be processed increase the size and the number of pixels of the images, these pixels may affect the inference calculation process of the visual model, resulting in errors in the inference results. Therefore, the features of these pixels need to be cropped before the inference output result.

[0083] For the convenience of subsequent cropping, after performing size padding processing on the images to be processed, the embodiments of the present invention construct a padding configuration file according to the size padding result. The size padding result of this padding configuration file records the size padding situation of the images to be processed, such as how many pixels are filled in the height and width, how many times the size is increased, etc. The size padding result of this padding configuration file is used to perform size cropping processing on the predicted feature map according to the size padding result before the visual model outputs the image inference result. The predicted feature map is the feature map used for mapping processing to obtain the image inference result, that is, the feature map before mapping processing through a fully connected layer neural network or a multi-layer perceptron.

[0084] For example, if the size of a certain image to be processed is 16×16 and the size doubles to 32×32 after size padding processing on the periphery of the image, then when the visual model performs inference and outputs the result, the periphery of the predicted feature map is cropped to half of the original size, and then the final image inference result is output using the predicted feature map after size cropping processing.

[0085] In the embodiments of the present invention, after the image to be processed is segmented, size filling processing is performed on the image to be processed, so that for images to be processed of various sizes, they can all be segmented according to a preset size, ensuring the integrity of the image. And through post - cropping processing, the influence of the filled part on the inference process is also eliminated, ensuring that the output image inference result is not affected.

[0086] In some embodiments, after the image to be inferred is input into the visual model for inference to obtain an image inference result, the image index corresponding to the image slice of the image to be inferred and the feature index corresponding to the feature sub - graph are also removed from the block list, and the storage space of the video memory block used to store the image slice and the feature sub - graph is released in the physical video memory.

[0087] Since there is more than one image to be processed that needs to be input into the visual model for inference, when a certain image to be processed is inferred, the image index corresponding to the image slice of the processed image to be processed and the feature index corresponding to the feature sub - graph need to be removed. For example Figure 3 in, the image index, feature index, and mapping relationship recorded in the physical block number column of the block list are all cleared, thereby updating the block list to facilitate storing the image index, feature index, and mapping relationship corresponding to the next image to be processed.

[0088] After the block list is updated, next, the storage space of the video memory block used to store the image slice and the feature sub - graph is released in the physical video memory. For example Figure 4 in, after the image to be processed is inferred, the storage spaces of the video memory blocks 1, 2, 3 for storing the image slice in the physical video memory and the video memory blocks 4, 5, 6 for storing the feature sub - graph are all released to facilitate processing and storing the next image to be processed.

[0089] In this way, for each image to be processed, after using the visual model for inference, the block list and the storage space are updated and released in the above - mentioned manner.

[0090] In the embodiments of the present invention, after the image to be processed is inferred, the block list and the physical video memory are updated and the storage space is released in a timely manner, which can reasonably and effectively utilize the physical video memory and avoid overloading the use of the physical video memory and affecting the inference efficiency of the visual model.

[0091] Next, the visual model inference device based on image segmentation provided by the present invention will be described. The visual model inference device based on image segmentation described below can be mutually referred to with the visual model inference method based on image segmentation described above.

[0092] As Figure 5As shown in the figure, the visual model inference device based on image segmentation specifically includes: a segmentation module 501, a storage module 502, an acquisition module 503, and an inference module 504. Specifically, the segmentation module 501 is used to segment the image to be processed according to a preset size to obtain multiple image slices, and record the image index of each image slice in the block list; the storage module 502 is used to divide the physical video memory storing the image to be processed into multiple video memory blocks, and store the image slices through the video memory blocks; the acquisition module 503 is used to, when the visual model performs inference, obtain the image slices from the video memory blocks according to the block list, and splice the image slices according to the image index to obtain the image to be inferred; the inference module 504 is used to input the image to be inferred into the visual model for inference to obtain the image inference result.

[0093] It should be noted that the beneficial effects of the visual model inference device based on image segmentation here correspond to those of the visual model inference method based on image segmentation in the above text. Therefore, the beneficial effects of the visual model inference device based on image segmentation will not be elaborated here.

[0094] Figure 6 An example of the physical structure diagram of an electronic device is shown in Figure 6 As shown in the figure, the electronic device may include: a processor 610 (processor), a communication interface 620 (Communications Interface), a memory 630 (memory), and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the visual model inference method based on image segmentation, and the method includes: segmenting the image to be processed according to a preset size to obtain multiple image slices, and recording the image index of each image slice in the block list; dividing the physical video memory storing the image to be processed into multiple video memory blocks, and storing the image slices through the video memory blocks; when the visual model performs inference, obtaining the image slices from the video memory blocks according to the block list, and splicing the image slices according to the image index to obtain the image to be inferred; inputting the image to be inferred into the visual model for inference to obtain the image inference result.

[0095] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0096] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the visual model inference method based on image segmentation provided by the above-mentioned various methods. The method includes: segmenting the image to be processed according to a preset size to obtain multiple image slices, and recording the image index of each image slice in a block list; dividing the physical video memory storing the image to be processed into multiple video memory blocks, and storing the image slices through the video memory blocks; when the visual model performs inference, obtaining the image slices from the video memory blocks according to the block list, and splicing the image slices according to the image index to obtain an image to be inferred; inputting the image to be inferred into the visual model for inference to obtain an image inference result.

[0097] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the visual model inference method based on image segmentation provided by the above-mentioned various methods. The method includes: segmenting the image to be processed according to a preset size to obtain multiple image slices, and recording the image index of each image slice in a block list; dividing the physical video memory storing the image to be processed into multiple video memory blocks, and storing the image slices through the video memory blocks; when the visual model performs inference, obtaining the image slices from the video memory blocks according to the block list, and splicing the image slices according to the image index to obtain an image to be inferred; inputting the image to be inferred into the visual model for inference to obtain an image inference result.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual model inference method based on image segmentation, characterized in that Including: Segment the image to be processed according to a preset size to obtain multiple image slices, and record the image index of each image slice in a block list; Divide the physical video memory storing the image to be processed into multiple video memory blocks, and store the image slices through the video memory blocks; When the vision model performs inference, obtain the image slices from the video memory blocks according to the block list, and splice the image slices according to the image index to obtain an image to be inferred; Input the image to be inferred into the vision model for inference to obtain an image inference result; The step of inputting the image to be inferred into the vision model for inference to obtain an image inference result includes: During the process of the vision model extracting the feature map of the image to be inferred, obtain the feature map extracted by each layer of the network, and segment the feature map to obtain multiple feature sub-maps; Record the feature index of the feature sub-maps in the block list, and store the feature sub-maps through the video memory blocks divided by the physical video memory; When the vision model performs inference, query the feature sub-maps from the block list according to the feature index; Splice the feature sub-maps according to the feature index to obtain a feature map to be processed; Perform a mapping process on the feature map to be processed output by the last layer of the network in the vision model to obtain an image inference result.

2. The visual model inference method based on image segmentation according to claim 1, wherein The step of storing the image slices through the video memory blocks includes: Combine the image slices into image blocks according to a preset quantity; Record the image blocks and the number of image slices included in the image blocks in the block list, and establish a mapping relationship between the image blocks and the video memory blocks in the block list; Store the image slices in the image blocks into the video memory blocks according to the mapping relationship.

3. The visual model inference method based on image segmentation according to claim 1, wherein The step of obtaining the image slices from the video memory blocks according to the block list includes: Query the mapping relationship between the image blocks and the video memory blocks from the block list according to the image index of the image slices; Determine the target video memory block corresponding to the storage of the image slices from the physical video memory according to the mapping relationship, and obtain the image slices with the corresponding image index according to the target video memory block.

4. The visual model inference method based on image segmentation according to claim 1, wherein Before segmenting the image to be processed according to the preset size, the method further includes: Perform size padding processing on the image to be processed, and construct a padding configuration file according to the size padding result; The padding configuration file is used to perform size cropping processing on the predicted feature map according to the size padding result before the vision model outputs the image inference result, and the predicted feature map is the feature map used for mapping processing to obtain the image inference result.

5. The visual model inference method based on image segmentation according to claim 1, characterized in that, After inputting the image to be inferred into the vision model for inference to obtain an image inference result, the method further includes: Remove the image index corresponding to the image slices of the image to be inferred and the feature index corresponding to the feature sub-maps from the block list; Release the storage space of the video memory blocks used to store the image slices and the feature sub-maps in the physical video memory.

6. A visual model inference device based on image segmentation, characterized in that, Including: A segmentation module, configured to segment the image to be processed according to a preset size, obtain a plurality of image slices, and record the image index of each image slice into a block list; A storage module, configured to divide the physical video memory storing the image to be processed into a plurality of video memory blocks, and store the image slices through the video memory blocks; An acquisition module, configured to, when the vision model performs inference, acquire the image slices from the video memory blocks according to the block list, and splice the image slices according to the image index to obtain an image to be inferred; An inference module, configured to input the image to be inferred into the vision model for inference to obtain an image inference result; The inputting the image to be inferred into the vision model for inference to obtain an image inference result includes: During the process of the vision model extracting the feature map of the image to be inferred, acquiring the feature maps extracted by each layer of the network, and segmenting the feature maps to obtain a plurality of feature sub-maps; Recording the feature index of the feature sub-maps into the block list, and storing the feature sub-maps through the video memory blocks divided by the physical video memory; When the vision model performs inference, querying the feature sub-maps from the block list according to the feature index; Performing splicing processing on the feature sub-maps according to the feature index to obtain a feature map to be processed; Performing mapping processing on the feature map to be processed output by the last layer of the network in the vision model to obtain an image inference result.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the vision model inference method based on image segmentation according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the vision model inference method based on image segmentation according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the vision model inference method based on image segmentation according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for improving super-division operation performance of AI computing chip

    CN115982418A

  • Image size adjusting method and device based on model reasoning

    CN116630145A