Method, apparatus, device, and medium for image processing or training of an image processing model
Through the image processing model of feature extraction, gridding and dynamic convolution kernel adjustment of shared bicycle images, the accuracy and efficiency of shared bicycle location and brand identification are solved, and efficient illegal parking detection is achieved.
Patent Information
- Application Number
- CN202310780968.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-28
AI Technical Summary
It is difficult for the existing technology to efficiently divide and identify the location and brand of shared bicycles, especially inadequate accuracy and efficiency in illegal parking detection.
By extracting the image feature, gridding, target category prediction and screening, image segmentation results are generated, and dynamic adjustment of dynamic convolution kernels is used to combine classification and positioning box regression to achieve accurate identification of shared bicycles and illegal parking detection.
It improves the accuracy and efficiency of image segmentation of shared bicycles, can identify bicycle brands and numbers at the same time, and enhances the accuracy and efficiency of illegal parking detection.
Smart Images

Figure CN117078997B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, specifically to the fields of smart cities, deep learning, and computer vision, and particularly to a method, apparatus, device, and medium for image processing or training of an image processing model. Background Art
[0002] Currently, as a public transportation means, shared bicycles carry the short-distance travel needs of citizens and play a crucial role in the daily urban traffic. Therefore, the regulatory need for shared bicycles has emerged. The standardized parking of shared bicycles not only beautifies the urban appearance but also improves the usage efficiency of shared bicycles. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and medium for image processing or training of an image processing model.
[0004] According to one aspect of the present disclosure, there is provided an image processing method, including:
[0005] Performing feature extraction on an image to be processed to obtain image features of the image to be processed, wherein the size of the image features of the image to be processed is smaller than the size of the image to be processed;
[0006] Meshing the image features of the image to be processed to obtain mesh features of at least one mesh to be processed;
[0007] Performing target class prediction on the mesh to be processed according to the mesh features of the mesh to be processed to obtain a class confidence of the mesh to be processed;
[0008] Screening the mesh features of each mesh to be processed according to the class confidence of each mesh to be processed to obtain screened features of the image to be processed;
[0009] Generating an image segmentation result of the image to be processed based on the screened features of the image to be processed and the image features of the image to be processed.
[0010] According to one aspect of the present disclosure, there is provided an image processing apparatus, including: The image processing model implements the image processing method as described in any embodiment of the present disclosure;
[0011] A feature extraction module, configured to perform feature extraction on an image to be processed to obtain image features of the image to be processed, wherein the size of the image features of the image to be processed is smaller than the size of the image to be processed;
[0012] A target classification module, configured to mesh the image features of the image to be processed to obtain mesh features of at least one mesh to be processed;
[0013] A target classification module, configured to perform target category prediction on the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid, and obtain the category confidence of the to-be-processed grid;
[0014] A target segmentation module, configured to screen the to-be-processed grid features of each to-be-processed grid according to the category confidence of each to-be-processed grid, and obtain the screened features of the to-be-processed image;
[0015] A target segmentation module, configured to generate an image segmentation result of the to-be-processed image based on the screened features of the to-be-processed image and the image features of the to-be-processed image.
[0016] According to one aspect of the present disclosure, there is provided a method for training an image processing model, including: the image processing model implements the image processing method described in any embodiment of the present disclosure;
[0017] Through the image processing model, feature extraction is performed on a sample image to obtain the image features of the sample image, and the size of the image features of the sample image is smaller than the size of the sample image;
[0018] Through the image processing model, the image features of the sample image are gridified to obtain the sample grid features of at least one sample grid;
[0019] Through the image processing model, target category prediction is performed on the sample grid according to the sample grid features of the sample grid, and the category confidence of the sample grid is obtained;
[0020] Through the image processing model, the sample grid features of each sample grid are screened according to the category confidence of each sample grid, and the screened features of the sample image are obtained;
[0021] Through the image processing model, a predicted segmentation result of the sample image is generated based on the screened features of the sample image and the image features of the sample image;
[0022] Through the image processing model, a first difference between the standard segmentation result of the sample image and the predicted segmentation result is calculated;
[0023] Through the image processing model, a second difference between the standard category of the sample image and the category confidence of each sample grid is calculated;
[0024] Through the image processing model, the model parameters of the image processing model are adjusted according to the first difference and the second difference.
[0025] According to one aspect of the present disclosure, there is provided a training device for an image processing model, including: the image processing model is configured as the image processing device according to any embodiment of the present disclosure;
[0026] The image processing model is used to extract features from a sample image to obtain the image features of the sample image, and the size of the image features of the sample image is smaller than the size of the sample image;
[0027] The image processing model is used to grid the image features of the sample image to obtain the sample grid features of at least one sample grid;
[0028] The image processing model is used to predict the target category of the sample grid according to the sample grid features of the sample grid to obtain the category confidence of the sample grid;
[0029] The image processing model is used to screen the sample grid features of each sample grid according to the category confidence of each sample grid to obtain the screened features of the sample image;
[0030] The image processing model is used to generate a predicted segmentation result of the sample image based on the screened features of the sample image and the image features of the sample image;
[0031] The image processing model is used to calculate a first difference between the standard segmentation result of the sample image and the predicted segmentation result;
[0032] The image processing model is used to calculate a second difference between the standard category of the sample image and the category confidence of each sample grid;
[0033] The image processing model is used to adjust the model parameters of the image processing model according to the first difference and the second difference.
[0034] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0035] At least one processor; and
[0036] A memory communicatively connected to the at least one processor; wherein,
[0037] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method or the training method of the image processing model according to any embodiment of the present disclosure.
[0038] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image processing method or the training method of the image processing model according to any embodiment of the present disclosure.
[0039] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the image processing method or the training method of the image processing model according to any embodiment of the present disclosure when executed by a processor.
[0040] The embodiments of the present disclosure can improve the accuracy and efficiency of target segmentation.
[0041] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0043] Figure 1 is a flowchart of an image processing method disclosed according to an embodiment of the present disclosure;
[0044] Figure 2 is a flowchart of another image processing method disclosed according to an embodiment of the present disclosure;
[0045] Figure 3 is a flowchart of another image processing method disclosed according to an embodiment of the present disclosure;
[0046] Figure 4 is a flowchart of a training method of an image processing model disclosed according to an embodiment of the present disclosure;
[0047] Figure 5 is a scene graph of an image processing method disclosed according to an embodiment of the present disclosure;
[0048] Figure 6 is a schematic diagram of an output result disclosed according to an embodiment of the present disclosure;
[0049] Figure 7 is a schematic structural diagram of an image processing device disclosed according to an embodiment of the present disclosure;
[0050] Figure 8 is a schematic structural diagram of a training device of an image processing model disclosed according to an embodiment of the present disclosure;
[0051] Figure 9It is a block diagram of an electronic device for an image processing method or a training method of an image processing model disclosed according to an embodiment of the present disclosure. Detailed implementation manners
[0052] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0053] Figure 1 It is a flowchart of an image processing method disclosed according to an embodiment of the present disclosure. This embodiment can be applied to the situation of segmenting and classifying target objects in an image. The method of this embodiment can be executed by an image processing device, which can be implemented in software and / or hardware and is specifically configured in an electronic device with a certain data operation ability. The electronic device can be a client device or a server device. Examples of client devices include mobile phones, tablet computers, vehicle-mounted terminals, and desktop computers.
[0054] S101. Extract features from the image to be processed to obtain the image features of the image to be processed, where the size of the image features of the image to be processed is smaller than the size of the image to be processed.
[0055] The image to be processed can refer to an image for identifying a target. Exemplarily, the target can be a shared non-motor vehicle, an animal, or a vehicle, etc. The shared non-motor vehicle can refer to a shared bicycle and a shared electric vehicle. Feature extraction can be achieved through convolution. The image features can be represented in the form of a matrix or a vector. The image to be processed can be regarded as a pixel matrix, and its size can refer to the length and width of the matrix. The size of the image features can refer to the length and width of the matrix describing the image features. Alternatively, the image features of the image to be processed can be understood as a feature map. The image features of the image to be processed are usually obtained by processing the image to be processed, and the size of the image features is smaller than the size of the image to be processed.
[0056] Optionally, extracting features from the image to be processed to obtain the image features of the image to be processed includes: performing multi-scale feature extraction on the image to be processed to obtain features of multiple sizes; and fusing the features of the multiple sizes to obtain the image features of the image to be processed.
[0057] Exemplarily, a backbone network can be used to extract features from the image to be processed, obtaining features of multiple sizes. A feature aggregation network is used to fuse the features of multiple sizes to obtain the image features of the image to be processed. Exemplarily, the backbone network can be Swin-Transformer, and the feature aggregation network can be FPN (Feature Pyramid Networks).
[0058] Among them, multi-scale feature extraction means that it is obtained through multiple stages (through operations such as convolution or linear transformation), and each stage processes to obtain a feature of one size, and different stages process to obtain features of different sizes.
[0059] By extracting multi-scale features, different-sized features can be fully utilized, fusing features rich in semantic information and features containing detailed information, enriching the content of the features. Thus, based on the rich features for classification and segmentation, the accuracy of object classification and segmentation can be improved.
[0060] S102. Grid the image features of the image to be processed to obtain the to-be-processed grid features of at least one to-be-processed grid.
[0061] Divide the image to be processed into multiple grids to obtain at least one to-be-processed grid. Gridding the image features can refer to dividing the image features to obtain the to-be-processed grid features corresponding to the to-be-processed grids. Exemplarily, divide the image to be processed into S*S to-be-processed grids, and determine the to-be-processed grid features of the to-be-processed grid according to the content corresponding to the to-be-processed grid in the image features. Exemplarily, the content corresponding to the to-be-processed grid in the image features can be directly determined as the to-be-processed grid features of the to-be-processed grid, or convolution can also be performed to obtain the to-be-processed grid features of the to-be-processed grid.
[0062] Optionally, the step of gridding the image features of the image to be processed to obtain the to-be-processed grid features of at least one to-be-processed grid includes: interpolating the image features of the image to be processed to obtain the interpolated features of the image to be processed; dividing the interpolated features of the image to be processed to obtain the to-be-processed grid features of at least one to-be-processed grid.
[0063] In fact, the size of the image features of the image to be processed is smaller than that of the image to be processed. Interpolate the image features to form interpolated features of the same size as the image to be processed. The size of the interpolated features is the same as that of the image to be processed. Perform grid division on the interpolated features, and use the content of the area corresponding to the position of the grid to be processed in the division result as the grid feature to be processed of the grid to be processed. In addition, the interpolated features can be further convolved, and the convolution result of the area corresponding to the position of the grid to be processed in the division result is used as the grid feature to be processed of the grid to be processed.
[0064] By converting the object classification and segmentation of the image to be processed into the classification and segmentation of grids at different positions, it can help detect small objects and improve the detection accuracy of objects.
[0065] S103. According to the grid feature to be processed of the grid to be processed, perform object category prediction on the grid to be processed to obtain the category confidence of the grid to be processed.
[0066] The object category prediction can be to predict whether there is at least one category of objects in the grid to be processed. If there is an object of a certain category in the grid to be processed, the probability of the grid to be processed for this category is relatively high; if there is no object of this category in the grid to be processed, the probability of the grid to be processed for this category is relatively low.
[0067] The category confidence is used to describe the reliability of the category prediction of the grid to be processed. Process the grid features to be processed of S*S grids to be processed, for example, through convolution processing, to obtain an object category prediction result of S*S*C, where C represents the category probability of the grid to be processed, and C can represent the category probability of the grid to be processed for at least one category. Each grid to be processed has a category probability of C. According to the category probability of the grid to be processed, determine the category confidence of the grid to be processed. Exemplarily, the highest category probability can be determined as the category confidence, or the mean value of the category probabilities can be taken and determined as the category confidence.
[0068] S104. Screen the grid features to be processed of each grid to be processed according to the category confidence of each grid to be processed to obtain the screened features of the image to be processed.
[0069] Screen out the grids to be processed whose category confidence meets the reliable conditions, and determine the grid features to be processed of the remaining grids to be processed as the screened features of the image to be processed. Exemplarily, the reliable condition is that the category confidence is greater than or equal to the preset confidence threshold, or the reliable condition is that the category confidence is less than the preset confidence threshold, or the reliable condition is that the category confidence is within the preset confidence range. The reliable condition can be used to detect whether there is an object in the image area corresponding to the grid to be processed and / or whether the category of the object is reliable, etc.
[0070] S105. Generate an image segmentation result for the image to be processed based on the screening features and the image features of the image to be processed.
[0071] The screening features and the image features can be used as the features of the image to be processed for image segmentation of the image to be processed to obtain an image segmentation result. The image segmentation result can be a mask image, that is, different pixel values are used to distinguish the target and non-target pixels in the image to be processed. The image segmentation result is obtained based on the screening features, specifically including the position information of the grid and the category information of the key grids after screening. Therefore, in the process of image segmentation, the screening features provide richer and more accurate position and category information to improve the accuracy of image segmentation. Moreover, the screened features can be post-processed to reduce the amount of data to be processed and improve the efficiency of image segmentation.
[0072] Optionally, the image processing method is implemented by a pre-trained image processing model.
[0073] The image processing model is a machine learning model. Exemplarily, the image processing model includes a feature extraction module, a target classification module, and a target segmentation module. The feature extraction module is used to extract features from the input image. The target classification module is used to classify the image according to the features. The target segmentation module is used to perform target segmentation on the image. Implementing the image processing method through the image processing module can improve the segmentation accuracy.
[0074] According to the technical solution of the present disclosure, by performing feature extraction on the image to be processed, image features are obtained, and the image features are gridified to form the to-be-processed grid features of multiple to-be-processed grids. Then, class prediction is performed on the to-be-processed grids, and the to-be-processed grid features are screened according to the class confidence level to obtain the screening features. The screening features and the image features are processed to generate an image segmentation result for the image to be processed. In the process of image segmentation, the position sensitivity and target sensitivity of the detection head are increased to improve the segmentation accuracy. At the same time, based on the processed data after screening, the amount of data to be processed is reduced, and the efficiency of image segmentation is improved.
[0075] Figure 2 It is a flowchart of another image processing method disclosed according to an embodiment of the present disclosure, which is further optimized and extended based on the above technical solution and can be combined with the above various optional embodiments. The step of screening the to-be-processed grid features of each to-be-processed grid according to the class confidence level of each to-be-processed grid to obtain the screening features of the to-be-processed image is specifically: according to the class confidence level of each to-be-processed grid, the target to-be-processed grids with targets are screened out; the to-be-processed grid features of each target to-be-processed grid are combined to generate the screening features of the to-be-processed image.
[0076] S201. Extract features from the image to be processed to obtain the image features of the image to be processed, where the size of the image features of the image to be processed is smaller than the size of the image to be processed.
[0077] S202. Grid the image features of the image to be processed to obtain the processed grid features of at least one processed grid.
[0078] S203. Based on the processed grid features of the processed grid, perform target class prediction on the processed grid to obtain the class confidence of the processed grid.
[0079] S204. According to the class confidence of each processed grid, filter out the target processed grids with targets.
[0080] The target processed grid with a target may refer to that there is a target in the image area corresponding to the target processed grid. A preset confidence threshold can be used to determine whether there is a target in the image area corresponding to the target processed grid. Exemplarily, a grid with a class confidence greater than or equal to the confidence threshold is a target processed grid. A grid with a class confidence less than the confidence threshold is not a target processed grid. Exemplarily, the preset confidence threshold can be 0.3.
[0081] S205. Combine the processed grid features of each target processed grid to generate the filtered features of the image to be processed.
[0082] Combining the processed grid features of the target processed grid means increasing the dimension and forming a multi-dimensional filtered feature from the processed grid features.
[0083] In a specific example, the processed grid features of multiple processed grids of the image to be processed are [Batch, n_dim, S, S], where Batch is the image to be processed, S*S is the number of grids of the processed grid, n_dim is the processed grid feature, and the processed grid features of S*S processed grids form n_dim. After filtering, the number of target processed grids is P, the grid features of P target processed grids are n_dim, and after combination, the filtered feature obtained is [Batch, ndim, P], indicating that for the image to be processed Batch, there are P target processed grids, and the corresponding P processed grid features form ndim.
[0084] S206. Based on the filtered features of the image to be processed and the image features of the image to be processed, generate the image segmentation result of the image to be processed.
[0085] Optionally, predicting the target category of the to-be-processed grid based on the to-be-processed grid features of the to-be-processed grid to obtain the category confidence of the to-be-processed grid includes: predicting the target category of the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid to obtain the probabilities of at least one category to which the to-be-processed grid belongs; selecting the highest probability from the probabilities of at least one category to which the to-be-processed grid belongs and determining it as the category confidence of the to-be-processed grid.
[0086] If there is a target of a certain category in the to-be-processed grid, it is determined that the probability of the to-be-processed grid belonging to this category is relatively high. If there is no target of this category in the to-be-processed grid, it is determined that the probability of the to-be-processed grid belonging to this category is relatively low. The number of categories can be multiple. Compare the probabilities of the to-be-processed grid belonging to this category and select the highest probability value as the category confidence of this to-be-processed grid.
[0087] Exemplarily, the category confidence of the to-be-processed grid can be [Batch, S, S, C], C(c1, c2, c3), where c1 is the probability of the first category, c2 is the probability of the second category, and c3 is the probability of the third category. If c3 is the largest, the category confidence of this to-be-processed grid is c3.
[0088] By predicting at least one category for the to-be-processed grid and determining the highest probability among the probabilities of each category as the category confidence of the to-be-processed grid, the detection of the category confidence can be simplified, and at the same time, the category confidence can be accurately characterized, thereby improving the screening effectiveness of the to-be-processed grid.
[0089] Optionally, generating the image segmentation result of the to-be-processed image based on the screening features and the image features of the to-be-processed image includes: performing convolution and size adjustment on the image features of the to-be-processed image to obtain the adjusted image features of the to-be-processed image, and the size of the adjusted image features of the to-be-processed image is the same as the size of the to-be-processed image; determining the screening features of the to-be-processed image as the convolution kernel, performing convolution on the adjusted image features of the to-be-processed image, generating the mask segmentation result in the to-be-processed image, and determining it as the image segmentation result of the to-be-processed image.
[0090] The size of the image features is smaller than the image size, and the size of the adjusted image features is the same as the to-be-processed image. Exemplarily, the image features are [Batch, A, H / 8, W / 8], where A is the number of categories, H is the height of the to-be-processed image, and W is the width of the to-be-processed image. The adjusted image features are [Batch, n_dim, H, W]. The size adjustment can be achieved by linear transformation. The convolution and size adjustment can be implemented by 1×1 convolution, group normalization, and rectified linear unit (ReLU), etc.
[0091] The screening feature is determined as a convolutional kernel, and the adjusted image features are convolved using the screening feature to obtain a mask image of the image to be processed, which is determined as the mask segmentation result of the image to be processed, that is, the image segmentation result. In fact, in a typical object recognition scenario, the distribution of objects in the image is relatively sparse. During the process of inferring an object once, only the features corresponding to a small part of the grids are effective features. By determining the screening feature as a convolutional kernel and dynamically determining the convolutional kernel according to the output of the convolutional learning method to replace the fixed convolutional kernel to participate in the convolutional calculation, and then generating the segmentation result, it is possible to increase the effective calculation, reduce the ineffective calculation, and achieve reducing memory consumption and improving calculation efficiency.
[0092] By using the screening feature as a dynamic convolutional kernel and performing dynamic convolution to achieve image segmentation, it is possible to reduce ineffective calculations, improve calculation efficiency, and effectively adjust the convolutional kernel, increasing the flexibility and adaptability of the convolution.
[0093] According to the technical solution of the present disclosure, by screening the target grids to be processed with targets, and combining the grid features to be processed of the target grids to be processed to form a screening feature, and performing segmentation according to the screening feature, it is possible to improve the accuracy of image segmentation, reduce the calculation amount of image segmentation, and improve the image segmentation efficiency.
[0094] Figure 3 It is a flowchart of another image processing method disclosed according to an embodiment of the present disclosure, which is further optimized and extended based on the above technical solution and can be combined with each of the above optional embodiments. The image processing method is optimized as: based on the screening feature of the image to be processed and the image features of the image to be processed, perform object bounding box detection on the image to be processed.
[0095] S301. Extract features from the image to be processed to obtain the image features of the image to be processed, and the size of the image features of the image to be processed is smaller than the size of the image to be processed.
[0096] S302. Grid the image features of the image to be processed to obtain the grid features to be processed of at least one grid to be processed.
[0097] S303. According to the grid features to be processed of the grid to be processed, perform object category prediction on the grid to be processed to obtain the category confidence of the grid to be processed; the category confidence includes the brand detection result of shared non-motor vehicles.
[0098] Optionally, the brand detection result of the shared non-motor vehicle includes yes or no; or the confidence detection result of the brand of the shared non-motor vehicle includes single vehicle or group of vehicles.
[0099] The brand detection results include the detection results of at least one brand. The detection result of each brand is yes or no, specifically represented by probability. Among them, yes means it is this brand, and no means it is not this brand. For another example, the detection result of each brand is single vehicle or group of vehicles, specifically represented by probability. Among them, single vehicle means there is one shared non-motor vehicle of this brand, and group of vehicles means there are multiple (two or more) shared non-motor vehicles of this brand.
[0100] By setting the confidence detection results of the brand, yes or no, or single vehicle or group of vehicles for another example, the diversity and flexibility of categories can be increased to adapt to diverse shared non-motor vehicle detection scenarios.
[0101] S304. Screen the to-be-processed grid features of each to-be-processed grid according to the category confidence of each to-be-processed grid to obtain the screened features of the to-be-processed image.
[0102] S305. Generate the image segmentation result of the to-be-processed image based on the screened features of the to-be-processed image and the image features of the to-be-processed image.
[0103] S306. Detect whether the shared non-motor vehicle is illegally parked according to the segmentation result of the shared non-motor vehicle in the to-be-processed image and the brand detection result of the shared non-motor vehicle.
[0104] In an application scenario, the target is a shared non-motor vehicle. The category confidence is the brand of the shared non-motor vehicle. In the to-be-processed image, the segmentation result and brand detection result of the target non-motor vehicle can be recognized. The detection rules for illegal parking of shared non-motor vehicles of different brands are different. Exemplarily, a shared non-motor vehicle of brand A is determined to be illegally parked outside a certain range of the subway station entrance, and a shared non-motor vehicle of brand B is determined to be illegally parked at the subway station entrance.
[0105] According to the brand detection result of a certain shared non-motor vehicle, determine the detection method for illegal parking of this shared non-motor vehicle, and determine the illegal parking area or standard parking area in the to-be-processed image according to this detection method. Based on the illegal parking area or standard parking area, detect the segmentation result of this shared motor vehicle to determine whether this shared non-motor vehicle is in the illegal parking area or standard parking area, so as to determine whether this shared non-motor vehicle is illegally parked.
[0106] Optionally, the image processing method further includes: performing target box detection on the to-be-processed image based on the screened features of the to-be-processed image and the image features of the to-be-processed image.
[0107] After obtaining the screened features of the to-be-processed image, target box detection can be performed on the to-be-processed image based on the screened features of the to-be-processed image and the image features of the to-be-processed image.
[0108] Object box detection can refer to locating an object in an image to be processed. Optionally, the image features of the image to be processed can be processed to obtain the position of the object box in the image to be processed. Exemplarily, the image features can be further convolved to obtain the position of the object box in the image to be processed. Optionally, the filtered features of the image to be processed can be processed to obtain the position of the object box in the image to be processed. Exemplarily, the filtered features can be further convolved to obtain the position of the object box in the image to be processed. Optionally, the object box detection results of the filtered features and the image features can be fused to obtain the object box.
[0109] By using the filtered features and the image features of the image to be processed to detect the object box in the image to be processed, the representativeness of key information can be increased, and then the positioning box can be detected, improving the accuracy of positioning box detection.
[0110] Optionally, the object box detection of the image to be processed based on the filtered features of the image to be processed and the image features of the image to be processed includes: detecting the coordinates of the object box in the image to be processed according to the filtered features of the image to be processed.
[0111] Convolve the filtered features of the image to be processed to obtain the coordinates of the object box in the image to be processed. The object box can be a rectangular box, and the coordinates can be the coordinates of the four vertices of the object box.
[0112] As in the previous example, the filtered feature is a matrix of Batch*n_dim*P. Convolve this matrix until a matrix of Batch*P*4 is obtained, representing the 4 vertex coordinates of the Pi-th object box in the Batch of the image to be processed.
[0113] By using the filtered features to be processed to detect the object box in the image to be processed, the accuracy of positioning box detection can be improved, and the efficiency of positioning box detection can be increased.
[0114] According to the technical solution of the present disclosure, by setting the application scenario to the detection of illegal parking of shared non-motor vehicles, the segmentation results of each shared non-motor vehicle and the brand detection results of shared non-motor vehicles in the image to be processed can be detected in parallel, and based on this, the illegal parking of each brand of shared non-motor vehicles can be detected, which can accelerate the efficiency of illegal parking detection and take into account the detection accuracy.
[0115] Figure 4FIG. 0 is a flowchart of another method for training an image processing model according to an embodiment of the present disclosure. This embodiment can be applied to the case of training an image processing model for segmenting and classifying target objects in an image. The method of this embodiment can be executed by an image processing device, which can be implemented in software and / or hardware and is specifically configured in an electronic device with certain data computing capabilities. The electronic device can be a client device or a server device. Examples of client devices include mobile phones, tablet computers, vehicle-mounted terminals, and desktop computers. The image processing model implements the image processing method described in any embodiment of the present disclosure.
[0116] S401. Extract features from the sample image through the image processing model to obtain the image features of the sample image, where the size of the image features of the sample image is smaller than the size of the sample image.
[0117] The sample image is used to train the image processing model, and the sample image and the image to be processed have the same size.
[0118] S402. Grid the image features of the sample image through the image processing model to obtain the sample grid features of at least one sample grid.
[0119] The sample is divided into multiple grids to obtain at least one sample grid. Gridding the image features may refer to dividing the image features to obtain the sample grid features corresponding to the sample grid.
[0120] S403. Through the image processing model, predict the target category of the sample grid according to the sample grid features of the sample grid to obtain the category confidence of the sample grid.
[0121] The target category prediction may be to predict whether there is at least one category of target in the sample grid. If there is a target of a certain category in the sample grid, the probability of the sample grid for this category is relatively high; if there is no target of this category in the sample grid, the probability of the sample grid for this category is relatively low.
[0122] S404. Through the image processing model, screen the sample grid features of each sample grid according to the category confidence of each sample grid to obtain the screened features of the sample image.
[0123] Screen out the sample grids whose category confidence meets the reliable condition, and determine the sample grid features of the retained sample grids as the screened features of the sample image. Exemplarily, the reliable condition is that the category confidence is greater than or equal to the preset confidence threshold, or the reliable condition is that the category confidence is less than the preset confidence threshold, or the reliable condition is that the category confidence is within the preset confidence range. The reliable condition is used to describe the category of the target in the image area corresponding to the sample grid.
[0124] S405. Generate a predicted segmentation result of the sample image through the image processing model based on the selected features and the image features of the sample image.
[0125] The selected features and the image features can be used as the features of the sample image for image segmentation of the sample image to obtain a predicted segmentation result. The predicted segmentation result can be a mask image, that is, different pixel values are used to distinguish the target and non-target pixels in the sample image. The predicted segmentation result is obtained based on the selected features, specifically including the position information of the grid and the category information of the selected key grids. Therefore, during the image segmentation process, the selected features provide richer and more accurate position and category information to improve the accuracy of image segmentation. Moreover, the selected features can be post-processed to reduce the amount of data processed and improve the efficiency of image segmentation.
[0126] S406. Calculate a first difference between the standard segmentation result and the predicted segmentation result of the sample image through the image processing model.
[0127] The standard segmentation result can be the correct segmentation result of the sample image. The first difference is used to describe the difference between the predicted segmentation result and the ground truth.
[0128] S407. Calculate a second difference between the standard category of the sample image and the category confidence of each sample grid through the image processing model.
[0129] The standard category can be the correct classification result of the sample image. The second difference is used to describe the difference between the category confidence and the ground truth.
[0130] S408. Adjust the model parameters of the image processing model according to the first difference and the second difference through the image processing model.
[0131] With the goal of narrowing the first difference and the second difference, adjust the model parameters of the image processing model. The value of the loss function can be calculated according to the first difference and the second difference. When the loss function converges, it is determined that the training of the image processing model is completed.
[0132] Optionally, the standard category of the sample image includes: the target categories existing in the area of the sample image corresponding to the sample grid, where the number of target categories is at least one.
[0133] The standard category is the category of the target existing in the area corresponding to the sample grid. There can be at least one target, and the categories of different targets can be the same or different. Exemplarily, the categories of the existing targets can be combined to form the standard category.
[0134] By determining the standard category of the sample image as at least one of the categories of the objects existing in the area corresponding to the sample grid, the detection granularity of the category can be increased, and the accuracy of category detection can be improved.
[0135] Optionally, calculating the second difference between the standard category of the sample image and the category confidence of the sample grid includes: calculating the binary cross-entropy loss value between the category confidence of the sample grid and at least one target category in the corresponding area of the sample image.
[0136] The difference between the category confidence and the standard category can be represented by the binary cross-entropy loss value (BinaryCross Entropy loss, BCE loss).
[0137] When there are multiple target categories, the binary cross-entropy loss value between the category confidence and only one of the multiple target categories can be calculated. Alternatively, the binary cross-entropy loss value between the category confidence and the multiple target categories can be calculated.
[0138] Exemplarily, the categories of the objects existing in the area corresponding to the sample grid include category a, category b, and category c. The binary cross-entropy loss value between the category confidence and (i1 = 1, i2 = 1, i3 = 1) can be calculated. i1 is category a, i2 is category b, and i3 is category c. For another example, only using category c to calculate the binary cross-entropy loss value is specifically to calculate the binary cross-entropy loss value between the category confidence and (i1 = 0, i2 = 0, i3 = 1). Select the categories in sequence, and the categories selected later cover the categories selected earlier. The last selected category is determined as the target category for calculating the binary cross-entropy loss value.
[0139] By diversely selecting the categories for calculating the binary cross-entropy loss value, the detection accuracy of the categories can be flexibly configured.
[0140] Optionally, the training method of the image processing model further includes: detecting the sample box of the sample image through the image processing model based on the screening features and the image features of the sample image; adjusting the model parameters of the image processing model by the image processing model according to the first difference and the second difference includes: calculating the third difference between the standard box and the sample box of the sample image through the image processing model; adjusting the model parameters of the image processing model by the image processing model according to the first difference, the second difference, and the third difference.
[0141] The image processing model is also used to detect the positioning box of the sample image at the same time. The standard box is the correct target box of the target in the sample image. The sample box is the predicted target box obtained by the image processing model processing the sample image. The third difference is used to describe the difference between the sample box and the standard box.
[0142] The GIOU (Generalized Intersection over Union) can be used to calculate the third difference, and the boundary of the sample box can be supervised based on the third difference, so as to assist the learning of the target segmentation module. In fact, when the image processing model is actually implemented, the process of sample box detection can be eliminated. Only during training, the image processing model can be assisted to learn the boundary of the box to improve the recognition accuracy of the segmentation boundary.
[0143] By detecting the sample box through the image processing model and calculating the third difference between the sample box and the corresponding standard box, the longitudinal and lateral boundaries of the object can be perceived, so as to assist the learning of the target segmentation module.
[0144] Optionally, the detecting the sample box of the sample image based on the screening feature of the sample image and the image feature of the sample image includes: detecting the coordinates of the target box in the sample image according to the screening feature of the sample image to obtain the sample box of the sample image.
[0145] By using the screening feature and the image feature of the sample image to detect the target box in the sample image, the representativeness of key information can be increased, and then the positioning box can be detected, the detection accuracy of the positioning box can be improved, and further the supervision intensity of the target boundary can be increased to improve the segmentation ability of the image processing model.
[0146] Optionally, the training method of the image processing model further includes: through the image processing model, adjusting the number of grids of the sample image including sample grids according to the number of target categories to reduce the number of target categories existing in the area corresponding to the sample grids in the sample image; the gridifying the image feature of the sample image through the image processing model includes: through the image processing model, gridifying the image feature of the sample image according to the adjusted number of grids.
[0147] For accurate detection, usually one sample grid only includes one target category. The sample image is divided into grids to obtain sample grids. For each sample grid, the number of target categories included in the sample grid is counted. Adjust the size of the sample grid so that one sample grid preferably has only one target category as much as possible. By adjusting the number of grids included in the sample image, the size of the sample grid is adjusted. The more the number, the smaller the size; the fewer the number, the larger the size.
[0148] By performing convolution and gridification on the image features of the sample image based on the adjusted number of grids, the image of the sample grid can be enhanced, redundant information and interference information can be reduced, and the accuracy of image segmentation and classification can be improved.
[0149] According to the technical solution of the present disclosure, by adopting the screening features and image features to be processed, the target box can be detected in the image to be processed, the representativeness of key information can be increased, and then the positioning box can be detected, and the accuracy of positioning box detection can be improved.
[0150] Figure 5 and 6 is a scene graph of an image processing method disclosed according to an embodiment of the present disclosure. The specific method is as follows:
[0151] The structure of the overall network is as Figure 5 shown. The image processing model adopts the form of multi-task joint training, and a total of three branches of classification (target classification module), segmentation (target segmentation module) and localization (coordinate) box regression are set. Among them, the coordinate box regression branch and the segmentation branch locate the boundary of the object and have a mutually promoting effect; while the classification branch and the segmentation branch are associated in the form of a dynamic convolution kernel to play a mutually promoting role.
[0152] The input of the image processing model is an RGB (red, green, blue) picture, which is subjected to feature extraction through a backbone network (such as Swin-Transformer), and then after fusing multi-scale features through a feature aggregation network (such as FPN), the result is output after passing through the above three task branches.
[0153] The image processing model adopts a dynamic convolution kernel design: the reason it is called a dynamic convolution kernel is that the result of the segmentation branch is not learned by a fixed convolution kernel, but the convolution kernel is dynamically determined according to the output of a convolution kernel learning process and participates in the convolution operation, thereby generating a segmentation result.
[0154] For the aggregated image features [Batch, C, H / 8, W / 8], the classification branch is processed. The convolution kernel prediction process and the classification branch share several layers (such as 4 layers) of convolution, so as to enhance the association through feature sharing. The shared convolution layer is used to obtain the grid features to be processed.
[0155] The classification branch may include: meshing the image features of the to-be-processed image to obtain to-be-processed grid features of at least one to-be-processed grid; predicting the target category of the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid to obtain the category confidence of the to-be-processed grid. The convolutional kernel prediction process may include: on the basis of the classification branch, screening the to-be-processed grid features of each to-be-processed grid according to the category confidence of each to-be-processed grid to obtain the screened features of the to-be-processed image.
[0156] The output shape of the classification branch is [Batch, S, S, C] (where C is the number of categories and S is the number of grids, meaning that the feature map is divided into S*S regions, such as Figure 6 . Extract features [Batch, ndim, P] (P corresponds to the number of grids) with category confidence higher than the preset confidence threshold from the screened features, that is, the screened features, and its shape is [Batch, n_dim, S, S] (n_dim is the feature dimension, which can be understood as predicting one feature for each small grid). Use the screened features as the dynamic convolutional kernel to perform convolution on the adjusted image features [Batch, n_dim, H, W] to obtain the output of the target segmentation module as [Batch, P, H, W].
[0157] During model training, if the center of a target gt (ground truth) falls within a certain grid, sample features [Batch, ndim, P], that is, the screened features, corresponding to the sample grids with the target will be extracted from the sample grids. Perform convolution with the corresponding output [Batch, n_dim, H, W] of the adjusted segmentation module, and the output shape is [Batch, P, H, W]. Calculate the dice loss between the predicted segmentation result of the output and the gt for supervision.
[0158] For the classification branch, the score corresponding to the category of gt in the grid is 1, and the rest are 0, and a [Batch, S, S, C] target (label) can be generated, and the BCE loss between the standard category and the category confidence of the classification branch is calculated for supervision.
[0159] For the coordinate box localization branch, convolution is performed on the image features alone or on the screened features, and the GIOU loss is used to calculate the difference between the sample box and the standard box of the gt for supervision, which can enable the model to better perceive the horizontal and vertical boundaries of the object, thereby assisting the learning of the segmentation branch.
[0160] When applied in practice, the coordinate regression branch can be removed. This branch only provides additional supervision information during training to facilitate the learning of the segmentation branch. Therefore, only the classification and segmentation branches are needed during output. When the model is inferring, if the class score of a certain grid among the (S*S, C) outputs is higher than the confidence threshold (default 0.3), the mask output corresponding to the index in the output of the segmentation branch will be obtained as the segmentation map. To support different brands and the judgment of illegal parking scales, C = N*2 can be set, where N is the number of brands. In addition, for each brand, the judgment of single or group can be set additionally.
[0161] In addition, the coordinate regression branch can also be trained separately and connected in series with the image processing model for joint training.
[0162] Existing solutions only involve the position detection of single bicycles or electric vehicles. However, to achieve precise supervision of shared bicycles, refined algorithm recognition is required. For example, bicycles of different brands need to be connected to the processing processes of different manufacturers, and different numbers of bicycles also correspond to illegal parking events of different scales. Therefore, information such as brand and quantity needs to be output while detecting the position.
[0163] The embodiments of the present disclosure enrich the product function dimension. In addition to the basic function of identifying the positions of shared bicycles and shared electric vehicles, the product functions of brand recognition and scale judgment are also added. In addition, due to the numerous functions, the joint training solution can enable these strongly related tasks to promote each other and improve the recognition effect.
[0164] According to the embodiments of the present disclosure, Figure 7 is the structural diagram of the image processing device in the embodiments of the present disclosure. The embodiments of the present disclosure are applicable to the situation of segmenting and classifying target objects in images. The device is implemented by software and / or hardware and is specifically configured in an electronic device with certain data operation capabilities.
[0165] Such as Figure 7 An image processing device 700 shown includes: a feature extraction module 701, a target classification module 702, and a target segmentation module 703. Among them,
[0166] The feature extraction module 701 is used to extract features from the image to be processed to obtain the image features of the image to be processed, and the size of the image features of the image to be processed is smaller than the size of the image to be processed;
[0167] The target classification module 702 is used to grid the image features of the image to be processed to obtain the to-be-processed grid features of at least one to-be-processed grid;
[0168] The target classification module 702 is configured to perform target category prediction on the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid, and obtain the category confidence of the to-be-processed grid;
[0169] The target segmentation module 703 is configured to screen the to-be-processed grid features of each of the to-be-processed grids according to the category confidence of each of the to-be-processed grids, and obtain the screened features of the to-be-processed image;
[0170] The target segmentation module 703 is configured to generate an image segmentation result of the to-be-processed image based on the screened features of the to-be-processed image and the image features of the to-be-processed image.
[0171] According to the technical solution of the present disclosure, by performing feature extraction on the to-be-processed image, image features are obtained, and the image features are gridded to form the to-be-processed grid features of multiple to-be-processed grids. Then, category prediction is performed on the to-be-processed grids, and the to-be-processed grid features are screened according to the category confidence to obtain the screened features. The screened features and the image features are processed to generate the image segmentation result of the to-be-processed image. During the image segmentation process, the position sensitivity and target sensitivity of the detection head are increased, the segmentation accuracy is improved, and at the same time, based on the screened data processing, the amount of data to be processed is reduced, and the image segmentation efficiency is improved.
[0172] Further, the target segmentation module includes: a grid screening unit configured to screen out the target to-be-processed grids with targets according to the category confidence of each of the to-be-processed grids; a feature screening unit configured to combine the to-be-processed grid features of each of the target to-be-processed grids to generate the screened features of the to-be-processed image.
[0173] Further, the region motion detection module includes: a category prediction unit configured to perform target category prediction on the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid, and obtain the probabilities of at least one category to which the to-be-processed grid belongs; a confidence determination unit configured to select the highest probability among the probabilities of at least one category to which the to-be-processed grid belongs and determine it as the category confidence of the to-be-processed grid.
[0174] Further, the target segmentation module includes: a size adjustment unit configured to perform convolution and size adjustment on the image features of the to-be-processed image to obtain the adjusted image features of the to-be-processed image, and the size of the adjusted image features of the to-be-processed image is the same as the size of the to-be-processed image; a dynamic convolution unit configured to determine the screened features of the to-be-processed image as the convolution kernel, and perform convolution on the adjusted image features of the to-be-processed image to generate a mask segmentation result in the to-be-processed image, which is determined as the image segmentation result of the to-be-processed image.
[0175] Further, the target classification module includes: a feature interpolation unit configured to interpolate the image features of the image to be processed to obtain the interpolated features of the image to be processed; and a grid division unit configured to perform grid division on the interpolated features of the image to be processed to obtain the to-be-processed grid features of at least one to-be-processed grid.
[0176] Further, the class confidence includes the brand detection result of the shared non-motor vehicle; the image processing device further includes: a violation detection module configured to, after generating the image segmentation result of the image to be processed, detect whether the shared non-motor vehicle is parked illegally according to the segmentation result of the shared non-motor vehicle in the image to be processed and the brand detection result of the shared non-motor vehicle. The image processing device further includes: a target box detection module configured to perform target box detection on the image to be processed based on the filtered features of the image to be processed and the image features of the image to be processed.
[0177] Further, the brand detection result of the shared non-motor vehicle includes yes or no; or the confidence detection result of the brand of the shared non-motor vehicle includes single bike or group of bikes.
[0178] Further, the image processing device further includes: a target box detection module configured to perform target box detection on the image to be processed based on the filtered features of the image to be processed and the image features of the image to be processed.
[0179] Further, the feature extraction module includes: a multi-scale feature extraction unit configured to perform multi-scale feature extraction on the image to be processed to obtain features of multiple sizes; and a multi-scale feature fusion unit configured to fuse the features of multiple sizes to obtain the image features of the image to be processed.
[0180] Further, the image processing device is configured in the image processing model.
[0181] The above image processing device can execute the image processing method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the image processing method.
[0182] According to an embodiment of the present disclosure, Figure 8 is a structural diagram of a training device of an image processing model in an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of training an image processing model for segmenting and classifying target objects in an image. The device is implemented by software and / or hardware, and is specifically configured in an electronic device with certain data operation capabilities. The image processing model is configured with the image processing device as described in any embodiment of the present disclosure.
[0183] Such as Figure 8A training device 800 for an image processing model as shown, includes: an image processing model 801. Among them,
[0184] The image processing model is used to extract features from a sample image to obtain the image features of the sample image, and the size of the image features of the sample image is smaller than the size of the sample image; grid the image features of the sample image to obtain sample grid features of at least one sample grid; perform target category prediction on the sample grid according to the sample grid features of the sample grid to obtain the category confidence of the sample grid; screen the sample grid features of each sample grid according to the category confidence of each sample grid to obtain the screened features of the sample image; generate a predicted segmentation result of the sample image based on the screened features of the sample image and the image features of the sample image; calculate a first difference between the standard segmentation result of the sample image and the predicted segmentation result; calculate a second difference between the standard category of the sample image and the category confidence of each sample grid; and adjust the model parameters of the image processing model according to the first difference and the second difference.
[0185] Further, the standard category of the sample image includes: the target categories existing in the area of the sample image corresponding to the sample grid, where the number of target categories is at least one.
[0186] Further, the image processing model is further used to: calculate the binary cross-entropy loss value between the category confidence of the sample grid and at least one target category in the corresponding area of the sample image.
[0187] Further, the image processing model is used to perform sample box detection on the sample image based on the screened features of the sample image and the image features of the sample image; the image processing model is used to calculate a third difference between the standard box of the sample image and the sample box; the image processing model is used to adjust the model parameters of the image processing model according to the first difference, the second difference and the third difference.
[0188] Further, the image processing model is used to detect the coordinates of the target box in the sample image according to the screened features of the sample image to obtain the sample box of the sample image.
[0189] Further, the image processing model is used to adjust the grid number of the sample image including the sample grid according to the number of target categories to reduce the number of target categories existing in the area of the sample image corresponding to the sample grid; the image processing model is used to grid the image features of the sample image according to the adjusted grid number.
[0190] The training device of the above image processing model can execute the training method of the image processing model provided in any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution of the training method of the image processing model.
[0191] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0192] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0193] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0194] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0195] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0196] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as an image processing method or a training method of an image processing model. For example, in some embodiments, the image processing method or the training method of the image processing model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image processing method or the training method of the image processing model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the image processing method or the training method of the image processing model in any other suitable manner (e.g., by means of firmware).
[0197] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0198] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0199] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0200] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0201] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN) blockchain network, and the Internet.
[0202] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system or a server combined with blockchain.
[0203] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has technologies at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0204] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0205] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitations are imposed herein.
[0206] The above specific embodiments do not constitute a limitation to the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An image processing method, comprising: Performing feature extraction on the image to be processed to obtain the image features of the image to be processed, wherein the size of the image features of the image to be processed is smaller than the size of the image to be processed; Meshing the image features of the image to be processed to obtain the processed grid features of at least one processed grid; Performing target category prediction on the processed grid according to the processed grid features of the processed grid to obtain the category confidence of the processed grid; Screening the processed grid features of each processed grid according to the category confidence of each processed grid to obtain the screened features of the image to be processed; Generating an image segmentation result of the image to be processed based on the screened features of the image to be processed and the image features of the image to be processed; The screening the processed grid features of each processed grid according to the category confidence of each processed grid to obtain the screened features of the image to be processed includes: Screening out the target processed grids with targets according to the category confidence of each processed grid; Combining the processed grid features of each target processed grid to generate the screened features of the image to be processed.
2. The method according to claim 1, wherein The performing target category prediction on the processed grid according to the processed grid features of the processed grid to obtain the category confidence of the processed grid includes: Performing target category prediction on the processed grid according to the processed grid features of the processed grid to obtain the probabilities of at least one category to which the processed grid belongs; Selecting the highest probability among the probabilities of at least one category to which the processed grid belongs and determining it as the category confidence of the processed grid.
3. The method according to claim 1, wherein The generating an image segmentation result of the image to be processed based on the screened features of the image to be processed and the image features of the image to be processed includes: Performing convolution and size adjustment on the image features of the image to be processed to obtain the adjusted image features of the image to be processed, wherein the size of the adjusted image features of the image to be processed is the same as the size of the image to be processed; Determining the screened features of the image to be processed as a convolution kernel, performing convolution on the adjusted image features of the image to be processed, and generating a mask segmentation result in the image to be processed, which is determined as the image segmentation result of the image to be processed.
4. The method according to claim 1, wherein, The meshing the image features of the image to be processed to obtain the processed grid features of at least one processed grid includes: Interpolating the image features of the image to be processed to obtain the interpolated features of the image to be processed; Performing grid division on the interpolated features of the image to be processed to obtain the processed grid features of at least one processed grid.
5. The method according to claim 1, wherein, The category confidence includes the brand detection result of shared non-motor vehicles; After generating the image segmentation result of the image to be processed, it further includes: Detecting whether the shared non-motor vehicle is parked illegally according to the segmentation result of the shared non-motor vehicle in the image to be processed and the brand detection result of the shared non-motor vehicle.
6. The method according to claim 5, wherein, The brand detection result of the shared non-motor vehicle includes yes or no; or the confidence detection result of the brand of the shared non-motor vehicle includes single vehicle or group of vehicles.
7. The method according to claim 1, further comprising: Performing target box detection on the to-be-processed image based on the screening features of the to-be-processed image and the image features of the to-be-processed image.
8. The method according to claim 1, wherein The performing feature extraction on the to-be-processed image to obtain the image features of the to-be-processed image includes: Performing multi-scale feature extraction on the to-be-processed image to obtain features of multiple sizes; Fusing the features of the multiple sizes to obtain the image features of the to-be-processed image.
9. The method according to claim 1, wherein The image processing method is implemented by a pre-trained image processing model.
10. A training method for an image processing model, comprising: The image processing model implements the image processing method according to any one of claims 1-9; Through the image processing model, performing feature extraction on a sample image to obtain the image features of the sample image, and the size of the image features of the sample image is smaller than the size of the sample image; Through the image processing model, gridifying the image features of the sample image to obtain sample grid features of at least one sample grid; Through the image processing model, performing target class prediction on the sample grid according to the sample grid features of the sample grid to obtain the class confidence of the sample grid; Through the image processing model, screening the sample grid features of each sample grid according to the class confidence of each sample grid to obtain the screening features of the sample image; Through the image processing model, generating a predicted segmentation result of the sample image based on the screening features of the sample image and the image features of the sample image; Through the image processing model, calculating a first difference between the standard segmentation result of the sample image and the predicted segmentation result; Through the image processing model, calculating a second difference between the standard class of the sample image and the class confidence of each sample grid; Through the image processing model, adjusting the model parameters of the image processing model according to the first difference and the second difference.
11. The method according to claim 10, wherein, The standard class of the sample image includes: the target classes existing in the area of the sample image corresponding to the sample grid, where the number of target classes is at least one.
12. The method according to claim 11, wherein, The calculating the second difference between the standard class of the sample image and the class confidence of the sample grid includes: Calculating the binary cross-entropy loss value between the class confidence of the sample grid and at least one target class in the corresponding area of the sample image.
13. The method according to claim 12, further comprising: Through the image processing model, performing sample box detection on the sample image based on the screening features of the sample image and the image features of the sample image; The adjusting the model parameters of the image processing model according to the first difference and the second difference through the image processing model includes: Through the image processing model, calculating a third difference between the standard box of the sample image and the sample box. Adjust the model parameters of the image processing model according to the first difference, the second difference, and the third difference through the image processing model.
14. The method according to claim 13, wherein The sample box detection of the sample image based on the screening features of the sample image and the image features of the sample image includes: Detect the coordinates of the target box in the sample image according to the screening features of the sample image to obtain the sample box of the sample image.
15. The method according to claim 10, further comprising: Adjust the number of grids of the sample image including sample grids according to the number of target categories through the image processing model to reduce the number of target categories existing in the area corresponding to the sample grids in the sample image; The gridification of the image features of the sample image through the image processing model includes: Gridify the image features of the sample image according to the adjusted number of grids through the image processing model.
16. An image processing device, comprising: A feature extraction module for extracting features of a to-be-processed image to obtain the image features of the to-be-processed image, wherein the size of the image features of the to-be-processed image is smaller than the size of the to-be-processed image; A target classification module for gridifying the image features of the to-be-processed image to obtain the to-be-processed grid features of at least one to-be-processed grid; The target classification module for predicting the target category of the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid to obtain the category confidence of the to-be-processed grid; A target segmentation module for screening the to-be-processed grid features of each to-be-processed grid according to the category confidence of each to-be-processed grid to obtain the screening features of the to-be-processed image; The target segmentation module for generating an image segmentation result of the to-be-processed image based on the screening features of the to-be-processed image and the image features of the to-be-processed image; The target segmentation module includes: A grid screening unit for screening out the target to-be-processed grids with targets according to the category confidence of each to-be-processed grid; A feature screening unit for combining the to-be-processed grid features of each target to-be-processed grid to generate the screening features of the to-be-processed image.
17. The apparatus according to claim 16, wherein, A regional motion detection module, including: A category prediction unit for predicting the target category of the to-be-processed grid according to the to-be-processed grid features of the to-be-processed grid to obtain the probability of at least one category to which the to-be-processed grid belongs; A confidence determination unit for selecting the highest probability from the probabilities of at least one category to which the to-be-processed grid belongs and determining it as the category confidence of the to-be-processed grid.
18. The apparatus according to claim 16, wherein The target segmentation module includes: A size adjustment unit for performing convolution and size adjustment on the image features of the to-be-processed image to obtain the adjusted image features of the to-be-processed image, wherein the size of the adjusted image features of the to-be-processed image is the same as the size of the to-be-processed image; A dynamic convolution unit, which is used to determine the filtered features of the image to be processed as a convolution kernel, perform convolution on the adjusted image features of the image to be processed, generate a mask segmentation result in the image to be processed, and determine it as the image segmentation result of the image to be processed.
19. The apparatus according to claim 16, wherein, The target classification module includes: A feature interpolation unit, which is used to interpolate the image features of the image to be processed to obtain the interpolated features of the image to be processed; A grid division unit, which is used to divide the interpolated features of the image to be processed into grids to obtain the processed grid features of at least one processed grid.
20. The apparatus according to claim 16, wherein, The category confidence includes the brand detection result of the shared non-motor vehicle; The device further includes: A violation detection module, which is used to detect whether the shared non-motor vehicle is parked illegally according to the segmentation result of the shared non-motor vehicle in the image to be processed and the brand detection result of the shared non-motor vehicle after generating the image segmentation result of the image to be processed; The device further includes: A target box detection module, which is used to perform target box detection on the image to be processed based on the filtered features of the image to be processed and the image features of the image to be processed.
21. The apparatus according to claim 20, wherein, The brand detection result of the shared non-motor vehicle includes yes or no; or the confidence detection result of the brand of the shared non-motor vehicle includes single vehicle or group of vehicles.
22. The device according to claim 16, further includes: A target box detection module, which is used to perform target box detection on the image to be processed based on the filtered features of the image to be processed and the image features of the image to be processed.
23. The apparatus according to claim 16, wherein, The feature extraction module includes: A multi-scale feature extraction unit, which is used to perform multi-scale feature extraction on the image to be processed to obtain features of multiple sizes; A multi-scale feature fusion unit, which is used to fuse the features of multiple sizes to obtain the image features of the image to be processed.
24. The apparatus according to claim 16, wherein The image processing device is configured in the image processing model.
25. A training device for an image processing model, comprising: The image processing model is configured with the image processing device according to any one of claims 16-24; The image processing model is used to extract features from a sample image to obtain the image features of the sample image, and the size of the image features of the sample image is smaller than the size of the sample image; The image processing model is used to grid the image features of the sample image to obtain the sample grid features of at least one sample grid; The image processing model is used to perform target category prediction on the sample grid according to the sample grid features of the sample grid to obtain the category confidence of the sample grid; The image processing model is used to screen the sample grid features of each sample grid according to the category confidence of each sample grid to obtain the filtered features of the sample image; The image processing model is used to generate a predicted segmentation result of the sample image based on the filtered features of the sample image and the image features of the sample image; The image processing model is used to calculate a first difference between the standard segmentation result and the predicted segmentation result of the sample image; The image processing model is used to calculate a second difference between the standard category of the sample image and the class confidence of each sample grid; The image processing model is used to adjust the model parameters of the image processing model according to the first difference and the second difference.
26. The apparatus according to claim 25, wherein, The standard category of the sample image includes: the target categories existing in the area of the sample image corresponding to the sample grid, where the number of target categories is at least one.
27. The apparatus according to claim 26, wherein, The image processing model is further used to: Calculate the binary cross-entropy loss value between the class confidence of the sample grid and at least one target category in the corresponding area of the sample image.
28. The apparatus according to claim 27, further comprising: The image processing model is used to perform sample box detection on the sample image based on the screening features of the sample image and the image features of the sample image; The image processing model is used to calculate a third difference between the standard box of the sample image and the sample box; The image processing model is used to adjust the model parameters of the image processing model according to the first difference, the second difference, and the third difference.
29. The apparatus according to claim 28, further comprising: The image processing model is used to detect the coordinates of the target box in the sample image according to the screening features of the sample image, and obtain the sample box of the sample image.
30. The apparatus according to claim 25, further comprising: The image processing model is used to adjust the number of grids of the sample image including the sample grid according to the number of target categories, so as to reduce the number of target categories existing in the area of the sample image corresponding to the sample grid; The image processing model is used to grid the image features of the sample image according to the adjusted number of grids.
31. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method according to any one of claims 1-9 or the training method of the image processing model according to any one of claims 10-15.
32. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the image processing method according to any one of claims 1-9 or the training method of the image processing model according to any one of claims 10-15.
33. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the image processing method according to any one of claims 1-9 or the training method of the image processing model according to any one of claims 10-15.
Citation Information
Patent Citations
Target detection and target detection model training method and device
CN114549926A
Strip noise detection model training method, strip noise detection method and device
CN116188897A