Image processing method and device, electronic equipment, storage medium and program product

Boundary box information is obtained through image segmentation, and image complexity is calculated, which solves the problem of high computational volume in the prior art, and realizes efficient and accurate image complexity evaluation.

CN120339314APending Publication Date: 2025-07-18ZHONGDIAN DATA IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419701.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing image complexity calculation methods are large in calculation, resulting in wasted computing resources.

Method used

The bounding box information is obtained through the image segmentation process, and the object distribution characteristics are determined based on the bounding box information, and the image complexity is then calculated.

Benefits of technology

It reduces the amount of calculation, improves the calculation efficiency and evaluation accuracy, and adapts to efficient image processing in different image complexity scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339314A_ABST
    Figure CN120339314A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of computer vision, and the method comprises the steps: carrying out the image segmentation processing of a to-be-processed image, and obtaining at least one bounding box and the bounding box information of the at least one bounding box; wherein the bounding box information is used for representing the position and size of an object corresponding to at least one bounding box in the to-be-processed image; based on the bounding box information of the at least one bounding box, determining an object distribution feature of the to-be-processed image, the object distribution feature being used for representing a spatial distribution state of an object in the to-be-processed image in the to-be-processed image; and determining the image complexity of the to-be-processed image corresponding to the object distribution feature based on the object distribution feature so as to perform preset image processing on the to-be-processed image. The subsequent complexity calculation is carried out by taking the object corresponding to the bounding box as a unit, so that the processed data volume is reduced, and the calculation efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to an image processing method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] Image complexity (IC) refers to the amount of information contained in an image and the complexity of the image structure. Image complexity is a common and important property in computer vision and has wide applications in fields such as image segmentation, image steganography, text detection, and image enhancement. For example, in an image segmentation task, image complexity can be used as prior information to help the algorithm better understand the image content and improve the accuracy of segmentation.

[0003] Most of the existing image complexity calculation methods are based on pixel-level analysis, such as methods based on image entropy, edge detection, etc. Although these methods can reflect the complexity of the image to a certain extent, pixel-level analysis methods need to process each pixel, resulting in a large amount of calculation and low efficiency, increasing the calculation burden of computer vision algorithms and thus causing waste of computing resources.

[0004] The above content is only used to assist in understanding the technical solution of the present application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of the present application is to provide an image processing method, apparatus, electronic device, storage medium, and program product, aiming to solve the technical problem that the calculation amount of the image complexity calculation method is large, resulting in waste of computing resources.

[0006] To achieve the above object, the present application proposes an image processing method, and the method includes:

[0007] Performing image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box; wherein, the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed;

[0008] Determining the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box, wherein the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed;

[0009] Determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature for performing preset image processing on the image to be processed.

[0010] In addition, to achieve the above object, the present application further provides an image processing apparatus, which includes:

[0011] An image segmentation module, configured to perform image segmentation processing on the image to be processed, to obtain at least one bounding box and bounding box information of the at least one bounding box, where the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed;

[0012] A feature determination module, configured to determine the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box, where the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed;

[0013] A complexity determination module, configured to determine the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, for performing preset image processing on the image to be processed.

[0014] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the image processing method as described above.

[0015] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the image processing method as described above are implemented.

[0016] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the image processing method as described above are implemented.

[0017] In the present application, the image to be processed is subjected to image segmentation processing to obtain at least one bounding box and bounding box information of the at least one bounding box, where the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed; the object distribution feature of the image to be processed is determined based on the bounding box information of the at least one bounding box, where the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed; the image complexity of the image to be processed corresponding to the object distribution feature is determined based on the object distribution feature, for performing preset image processing on the image to be processed.

[0018] This application directly obtains bounding box information through an image segmentation model, determines the object distribution characteristics in units of bounding boxes, and the object distribution characteristics reflect the spatial distribution state of objects in the image. That is, this application actually determines the image complexity in units of the objects corresponding to the bounding boxes. For example, when objects are close to each other, overlapping, or occluding, the positional relationship of the bounding boxes is complex and the image complexity is high; when they are separated and loosely distributed, the complexity is low. Another example is that when there are many objects in the image and they are densely distributed, the number of bounding boxes is large and dense, and the image complexity is high; when the number of objects is small and they are sparsely distributed, the number of bounding boxes is small and sparse, and the complexity is low.

[0019] Compared with the pixel-level analysis method that needs to process and calculate each pixel in the image, the amount of data is huge, resulting in a huge amount of calculation. In this application, the subsequent complexity calculation is carried out in units of the objects corresponding to the bounding boxes, reducing the amount of data to be processed, thereby significantly improving the calculation efficiency.

[0020] In addition, compared with the pixel-based analysis that is easily interfered by background pixels and it is difficult to accurately judge the distribution of objects, thus affecting the evaluation accuracy of image complexity. In this application, the object distribution characteristics are determined based on the bounding box information, and various factors such as the position, size of the objects and their spatial distribution relationship are considered in the evaluation process, making the calculated image complexity more in line with the actual complexity of the image, thereby improving the evaluation accuracy of image complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the image processing method of this application;

[0024] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the image processing method of this application;

[0025] Figure 3 It is a schematic flowchart of the brief image processing method provided for an embodiment of this application;

[0026] Figure 4 It is a schematic diagram of the processing result of image segmentation provided for an embodiment of this application;

[0027] Figure 5 It is a schematic diagram of the module structure of the image processing device according to an embodiment of the present application;

[0028] Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the image processing method according to an embodiment of the present application. Specific embodiments

[0029] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0030] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the specification drawings and specific embodiments.

[0031] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. The following takes an electronic device as an example to illustrate this embodiment and the following embodiments.

[0032] Based on this, an embodiment of the present application provides an image processing method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the image processing method of the present application.

[0033] In this embodiment, the image processing method includes steps S10 to S30:

[0034] Step S10, performing image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box, where the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed;

[0035] The image to be processed is image data that needs to be subjected to specific image processing operations, and specifically can come from various different data sources, such as photos taken by a camera, a certain frame in a video, satellite remote sensing images, etc. In this embodiment, the image is input into a preset image segmentation model, and the preset image segmentation model processes the input image to be processed and outputs at least one bounding box and its related bounding box information. It should be noted that the bounding box can be a single bounding box used to frame the approximate range of an object or area in the image. The bounding box information is used to characterize the position and size of the object within the bounding box in the image to be processed, and usually includes the upper left coordinates (x, y), width w, and height h of the bounding box. Through the bounding box information, the specific position of the bounding box in the image and the approximate range of the object it contains can be accurately determined.

[0036] In this embodiment, by identifying objects or areas in the image to be processed and outputting at least one bounding box and bounding box information corresponding to each bounding box, basic data is provided for subsequent analysis of image complexity. Through the bounding box and its information, objects in the image can be intuitively located and quantified, thereby improving processing efficiency and targeting.

[0037] Step S20, determining object distribution features of the image to be processed based on the bounding box information of the at least one bounding box, wherein the object distribution features are used to characterize a spatial distribution state of the objects in the image to be processed in the image to be processed;

[0038] The object distribution feature is used to characterize the spatial distribution state of objects in the image to be processed. The specific feature indicators can be set according to actual needs. For example, the object distribution feature may include feature indicators such as the number of objects, the size of objects, and the position distribution of objects. Among them, the number of objects represents the total number of objects in the image, reflects the density of objects in the image, and can be reflected by the number of bounding boxes; the size of objects can be reflected by the area of bounding boxes; the position distribution of objects reflects the specific position of objects in the image, and their relative positional relationship, which can be reflected by the positional relationship between bounding boxes. It can be understood that in the image to be processed, the more objects there are, the higher the image complexity is usually; the larger the area of the object, the higher the image complexity may be; the higher the overlap of the object position distribution, the higher the image complexity may be.

[0039] In this embodiment, the spatial distribution state between the bounding boxes is determined by the bounding box information, and the spatial distribution state between the bounding boxes is determined as the object distribution state. The specific process can be determined according to the specific characteristic index of the object distribution feature, which is not limited here. For example, when the object distribution feature includes the number of objects, the number of bounding boxes is accumulated and the number of bounding boxes is determined as the number of objects.

[0040] Step S30: determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, so as to perform preset image processing on the image to be processed.

[0041] Based on the object distribution characteristics, the image complexity of the image to be processed is determined by statistics or machine learning. For example, the index values of each feature index of the object distribution characteristics can be weighted, averaged, etc., and the image complexity can be quantified according to the statistical processing results. For another example, a machine learning model for complexity calculation can be established, and the index values of each feature index of the object distribution characteristics can be input into the machine learning model to output the image complexity.

[0042] In this embodiment, image processing is performed on the image to be processed based on the image complexity. If the image complexity is greater than a preset complexity threshold, a preset high-precision image processing algorithm is used to perform image processing on the image to be processed; if the image complexity is less than or equal to the complexity threshold, a preset low-precision image processing algorithm is used to perform image processing on the image to be processed. Among them, high-precision image processing algorithms are those that can process images more meticulously and accurately. High-precision algorithms often take into account more detailed information in the image and can provide higher-quality results during the processing, such as more precise edge detection, more delicate image enhancement effects, more accurate target recognition and segmentation, etc.; in contrast to high-precision algorithms, low-precision algorithms handle details relatively simply during image processing, have lower computational complexity, faster execution speed, but the quality of the processing results may be inferior to high-precision algorithms. It can be understood that in the image processing task, adjusting the precision of the image processing algorithm according to the image complexity. In high-complexity images, using high-precision algorithms can ensure that targets in complex scenes are accurately detected and recognized; while in low-complexity images, using low-precision algorithms can quickly complete the processing task, saving computational resources, thereby ensuring improved processing efficiency while meeting application requirements.

[0043] In this embodiment, the specific image processing process is not limited here and can be set according to the actual image processing scenario and image processing task. Exemplarily, in the image stitching scenario, if the image complexity is high, a more precise feature matching algorithm is selected or the image fusion step is increased; in the target detection scenario, if the complexity is high, the sliding window density is increased or a multi-scale detection method is used; in the robot navigation and obstacle avoidance scenario, when the complexity is high, the robot plans the path more carefully, optimizes the sensor usage and fusion strategy; in video surveillance, more computational resources are allocated to high-complexity regions for target detection and tracking; in the image generation and editing scenario, the generation model parameters are adjusted or the editing quality is evaluated according to the complexity. Exemplarily, for the target detection scenario, if the image complexity is higher than the complexity threshold (such as 0.8), a high-precision detection model (Faster R-CNN) is called and the candidate box density is increased; if the image complexity is lower than the threshold, a lightweight model (YOLO Nano) is called and the detection frequency is reduced.

[0044] In a feasible implementation manner, the step S10: performing image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box includes:

[0045] Step S101, inputting the image to be processed into a preset image segmentation model to obtain at least one bounding box and the bounding box information of the at least one bounding box.

[0046] In this embodiment, an image segmentation model is pre-trained. The image segmentation model uses training images as input data and is trained with training images with bounding boxes as label data. The bounding boxes are used to label objects in the training images. The specific training process will not be elaborated here. The image segmentation model is a deep learning model whose main function is to segment and identify different objects or regions in an image. The model assigns a class label to each pixel in the image, thereby dividing the image into multiple different parts. In this implementation, there is no limitation on the specific image segmentation model. For example, it can be a U-Net (Convolutional Networks for Biomedical Image Segmentation), a Mask R-CNN (Mask Region-based Convolutional Neural Network), a SAM (Segment Anything Model), etc., which can be set according to actual needs.

[0047] In a feasible embodiment, the image segmentation model is a SAM model. The SAM model includes an image encoder, a prompt encoder, and a decoder. The step S101: Perform image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box, including:

[0048] Step S1011, input the image to be processed into the image encoder to obtain the first output data;

[0049] The image encoder is a key component in the SAM model, responsible for converting the input image data into a feature representation. Usually based on deep learning architectures such as CNN (Convolutional Neural Network) or ViT (Vision Transformer, the Transformer architecture applied to models in the field of computer vision), the image encoder extracts high-level features from the image and outputs feature maps, which contain the semantic information and spatial information of the image. The prompt encoder is a component in the SAM model used to process the user input prompt information, capable of converting various prompt information provided by the user (such as points, boxes, text, etc.) into a feature representation that matches the image features. The output of the prompt encoder is usually a set of feature vectors, which are input into the decoder and combined with the image features to generate a segmentation mask. The decoder is a component in the SAM model responsible for generating the segmentation mask. It receives the outputs of the image encoder and the prompt encoder, fuses this feature information, and generates the final segmentation mask. The decoder is usually based on a convolutional neural network or other generative models and can generate a high-quality segmentation mask according to the input feature information.

[0050] In this embodiment, the image to be processed is input into the image encoder, and the image encoder extracts features from the image and outputs a set of feature maps containing the semantic information and spatial information of the image, and the feature maps are determined as the first output data.

[0051] Step S1012: Receive the user prompt information and input the user prompt information into the prompt encoder to obtain the second output data;

[0052] The user prompt information is auxiliary information provided by the user when using the SAM model, used to guide the model to generate a more accurate segmentation result. The prompt information can be in various forms, such as points, boxes, text, etc. For example, the prompt information can be a certain position in the image clicked by the user or a bounding box drawn, or it can be a text description, such as "segment all cats", which is not limited here. The acquisition method of the user prompt information is not limited here. For example, it can accept the user prompt information manually input by the user, or it can be the automatic detection result output by the automatic detection system as the user prompt information, such as the bounding box generated by the object detection algorithm, which is not limited here.

[0053] Receive the user prompt information and input it into the prompt encoder. The prompt encoder encodes the prompt information, converts the user prompt information into a feature vector that matches the image features, and determines the feature vector as the second output data.

[0054] Step S1013: Input the first output data and the second output data into the decoder to obtain the segmentation mask;

[0055] The segmentation mask is the output generated by the SAM model, which is used to represent which object or region each pixel in the image belongs to. The segmentation mask is usually a binary image, where the value of each pixel indicates whether the pixel belongs to a specific object or region. The first output data and the second output data are input into the decoder, and the decoder combines this feature information to generate the segmentation mask.

[0056] In step S1014, traverse each of the segmentation masks, determine the contour coordinates of the segmentation region corresponding to the segmentation mask in the image to be processed, generate a bounding box based on the contour coordinates, and determine the bounding box information of the bounding box based on the contour coordinates.

[0057] The mask generated by the SAM model (i.e., the segmentation mask) is usually a list, and each element in the list corresponds to the mask information of a segmentation region, which is stored in dictionary form. Traverse each segmentation mask in this list. For each mask, use a function of an image processing library (such as OpenCV) to find the contour of the corresponding connected region (i.e., the segmentation region) in its binary image. The found contour is composed of a series of points, and each point has its coordinates in the original image (i.e., the contour coordinates), thus realizing the determination of the contour coordinates of the segmentation region corresponding to the segmentation mask in the image to be processed. That is, the segmentation region corresponding to the segmentation mask refers to the set of pixels in the segmentation mask that are marked as belonging to a certain object or region. The segmentation region can be determined by the pixel values in the segmentation mask and is usually represented as one or more connected regions; the contour coordinates refer to the boundary coordinates of the segmentation region, which are used to represent the shape and position of the segmentation region. The contour coordinates can be extracted from the segmentation mask through image processing algorithms and are usually represented as a series of point coordinates.

[0058] Based on the contour coordinates of the segmentation region, calculate one or more bounding boxes, which are used to represent the position and size of the segmentation region. The bounding box is usually determined by four parameters: the coordinates (x, y) of the upper left corner, the width (w), and the height (h), that is, the bounding box information. It can be understood that by generating the bounding box and the bounding box information, it is possible to more conveniently represent and process the segmentation region, providing a basis for subsequent tasks such as image complexity evaluation.

[0059] In this embodiment, by performing image segmentation on the image to be processed, at least one bounding box and the bounding box information of the at least one bounding box are obtained, where the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed; based on the bounding box information of the at least one bounding box, the object distribution feature of the image to be processed is determined, where the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed; based on the object distribution feature, the image complexity of the image to be processed corresponding to the object distribution feature is determined, so as to perform a preset image processing on the image to be processed.

[0060] Compared with the pixel-level analysis method that needs to process and calculate each pixel in the image, the data volume is huge, resulting in a huge amount of calculation. In this application, the subsequent complexity calculation is carried out in units of the objects corresponding to the bounding boxes, reducing the amount of data to be processed, thus significantly improving the calculation efficiency.

[0061] In addition, compared with the pixel-based analysis that is easily affected by background pixels and it is difficult to accurately judge the object distribution, thus affecting the evaluation accuracy of the image complexity. In this application, the object distribution feature is determined based on the bounding box information. During the evaluation process, various factors such as the position, size of the objects and their spatial distribution relationship are considered, making the calculated image complexity more in line with the actual complexity of the image, thereby improving the evaluation accuracy of the image complexity.

[0062] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 2 , the step S20: determining the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box includes:

[0063] Step S201, determining the number of bounding boxes of the at least one bounding box, and determining the bounding box area of the at least one bounding box based on the bounding box information;

[0064] Counting the number of bounding boxes generated by the image segmentation model to obtain the number of bounding boxes, and calculating the total area of each bounding box according to the coordinate and size information of each bounding box to obtain the bounding box area. The number of bounding boxes can intuitively reflect the approximate number of objects in the image, and the bounding box area can reflect the scale of the objects. Generally speaking, the more the number of bounding boxes, the more the number of objects, the higher the image complexity, and the larger the bounding box area, the larger the object area, the higher the image complexity.

[0065] In this embodiment, the process of determining the area of the bounding box can be as follows: Detect whether there is overlap between each bounding box. If there is no overlap between each bounding box, for each bounding box, according to its width w and height h, use the formula area S = w * h to calculate the area of the bounding box, traverse each bounding box in the list, and determine the sum of the areas of each bounding box as the area of the bounding box; if there is overlap between each bounding box, calculate the area of the overlapping part, and calculate the sum of the areas of each bounding box, and obtain the area of the bounding box by subtracting the area of the overlapping part from the sum of the areas of each bounding box. Here, the method of calculating the area of the overlapping part is not limited. For example, in a feasible implementation, for every two bounding boxes, the overlapping area can be determined by comparing the coordinate ranges. Assume that the coordinates of two bounding boxes are (x1, y1, w1, h1) and (x2, y2, w2, h2) respectively. The upper left corner coordinates of the overlapping area are (max(x1, x2), max(y1, y2)), and the lower right corner coordinates are (min(x1 + w1, x2 + w2), min(y1 + h1, y2 + h2)). Then, according to these two coordinates, calculate the width and height of the overlapping area, and further obtain the overlapping area; in another feasible implementation, it can be to use a geometric calculation library such as the Shapely library, convert each bounding box into a polygon object, and then calculate the intersection of the two polygons through the intersection method to obtain the overlapping area.

[0066] Step S202: Determine the number of times the bounding box overlaps based on the bounding box information, where the number of times the bounding box overlaps is used to represent the number of times of overlap between any two bounding boxes in the at least one bounding box.

[0067] The number of times the bounding box overlaps is used to represent the number of times of overlap between any two bounding boxes in the at least one bounding box. In a set of bounding boxes, if there is an intersection between the areas of two bounding boxes, it is considered that these two bounding boxes overlap. By counting the overlapping situations between all pairs of bounding boxes, the number of times the bounding box overlaps can be obtained. The number of times the bounding box overlaps can reflect the complexity of the spatial distribution of objects in the image. Generally speaking, the more times the bounding box overlaps, the higher the degree of overlap between objects and the higher the complexity of the image.

[0068] Step S203: Determine the number of bounding boxes, the area of the bounding box, and the number of times the bounding box overlaps as the object distribution characteristics of the image to be processed.

[0069] It can be understood that the area of the bounding box reflects the size of each object in the image. By calculating the total area of all bounding boxes, the size distribution of the objects in the image can be determined. If there are a large number of large-area objects in the image, it indicates that the image usually contains more texture and details, and the complexity of the image is relatively high. The number of bounding boxes reflects the number of objects in the image. If there are a large number of objects in the image, it indicates that the content of the image is relatively complex, and the complexity of the image is relatively high. The number of times the bounding boxes overlap reflects the overlapping situation of the objects in the image. Overlapping objects will increase the visual complexity of the image, and the complexity of the image is relatively high. In this embodiment, by comprehensively determining the object distribution characteristics through multiple key indicators, it can more comprehensively and accurately reflect the distribution state of the objects in the image, provide richer and more accurate information for subsequent calculation of the image complexity, and thus process the image more reasonably, improving the accuracy and effectiveness of image processing.

[0070] In a feasible embodiment, the number of the at least one bounding box is greater than one; the step S202: determining the number of times the at least one bounding box overlaps based on the bounding box information includes:

[0071] Step S2021, dividing each bounding box into multiple groups of bounding box groups, where each group of bounding box groups includes two of the bounding boxes and each group of bounding box groups does not repeat;

[0072] All the bounding boxes are divided into multiple groups of bounding box groups, each group contains two bounding boxes, and each group of bounding box groups does not repeat. By dividing the bounding boxes into multiple groups of bounding box groups, it can be ensured that each pair of bounding boxes is only calculated once, avoiding repeated calculation and improving the calculation efficiency.

[0073] The specific grouping process is not limited here and can be set according to actual needs. Exemplarily, in a feasible implementation manner, nested loops can be used to achieve grouping. Initialize an empty list pairs to store the generated combinations. Use two nested loops to traverse the list. Among them, the outer loop starts from the first element and ends at the penultimate element, and the inner loop starts from the next element of the outer loop and ends at the last element; form a group of bounding boxes with the elements corresponding to the indexes of each pair of the outer loop and the inner loop and add them to the pairs list.

[0074] Step S2022, determining the overlapping box groups from the multiple groups of bounding box groups based on the bounding box information, and determining the number of the overlapping box groups in the multiple groups of bounding box groups as the number of times each of the bounding boxes overlaps, where the overlapping box group is a group of bounding box groups in which there is an overlapping area between the two bounding boxes in the group.

[0075] It should be noted that the overlapping bounding box group is a bounding box group in which there is an overlapping area between two bounding boxes within the group. Based on the bounding box information, it is determined whether each group of bounding box groups is an overlapping bounding box group, and then the number of overlapping bounding box groups among multiple groups of the bounding box groups is determined as the box overlapping times of each of the bounding boxes. In this embodiment, the method for determining whether each group of bounding box groups is an overlapping bounding box group is not limited herein. For example, based on the bounding box information, it can be determined whether there is an intersection over union (IoU) between two bounding boxes within the group or whether the IoU is greater than a threshold to determine whether it is an overlapping bounding box group. Another example is that based on the bounding box information, it can be determined whether the coordinates of two bounding boxes within the group overlap with each other to determine whether it is an overlapping bounding box group. Specifically, it can be set according to actual requirements.

[0076] It can be understood that if the number of bounding boxes of the at least one bounding box is less than or equal to one, then the box overlapping times is determined to be zero.

[0077] In a feasible embodiment, the bounding box group includes a first bounding box and a second bounding box; the step S2022: determining the overlapping bounding box group from multiple groups of the bounding box groups based on the bounding box information includes:

[0078] Step S20221, traverse each group of the bounding box groups, and based on the first bounding box information of the first bounding box, determine the first left abscissa value of the left corner point among each first boundary corner point, and determine the first right abscissa value of the right corner point among each of the first boundary corner points, and determine the first top ordinate value of the top corner point among each of the first boundary corner points, and determine the first bottom ordinate value of the bottom corner point among each of the first boundary corner points, where the first boundary corner point is each corner point of the first bounding box;

[0079] One of the bounding boxes in the bounding box group is called the first bounding box, and the other bounding box is called the second bounding box. It can be understood that the first bounding box and the second bounding box here can both be any bounding box in the bounding box group and do not constitute a sequential limitation.

[0080] Each corner point of the first bounding box is called a first boundary corner point, including the upper left corner point, the lower left corner point, the upper left corner point, and the lower left corner point. Among them, the first boundary corner points are divided into left corner points and right corner points in the horizontal direction. The left corner points are the upper left corner point and the lower left corner point, and the right corner points are the upper left corner point and the lower left corner point. The first boundary corner points are divided into top corner points and bottom corner points in the vertical direction. The top corner points are the upper left corner point and the upper right corner point, and the bottom corner points are the lower left corner point and the lower right corner point. The first left abscissa value is the abscissa value of the left corner point of the first bounding box, that is, the abscissa values of the upper left corner point and the lower left corner point. The first right abscissa value is the abscissa value of the right corner point of the first bounding box, that is, the abscissa values of the upper right corner point and the lower right corner point. The first top ordinate value is the ordinate value of the top corner point of the first bounding box, that is, the ordinate values of the upper left corner point and the upper right corner point. The first bottom ordinate value is the ordinate value of the bottom corner point of the first bounding box, that is, the ordinate values of the lower left corner point and the lower right corner point.

[0081] Traverse all bounding box groups. For the first bounding box in each bounding box group, extract the coordinate information of each first boundary corner point, including the abscissa value of the left corner point, the abscissa value of the right corner point, the ordinate value of the top corner point, and the ordinate value of the bottom corner point. Exemplarily, rect1 and rect2 represent the first bounding box and the second bounding box respectively. The bounding box information consists of four values, which are the upper left coordinates (x, y) of the bounding box and the width (w) and height (h) of the bounding box. That is, the bounding box information of rect1 is (x1, y1) and the width w1 and height h1. Then the first left abscissa value of rect1 is x1, the first right abscissa value of rect1 is x1, x1 + w1, the first top ordinate value of rect1 is y1, and the first bottom ordinate value of rect1 is y1 + h1.

[0082] Step S20222: Based on the second bounding box information of the second bounding box, determine the second left abscissa value of the left corner point among each second boundary corner point, and determine the second right abscissa value of the right corner point among each second boundary corner point, and determine the second top ordinate value of the top corner point among each second boundary corner point, and determine the second bottom ordinate value of the bottom corner point among each second boundary corner point, where the second boundary corner point is each corner point of the second bounding box;

[0083] Each corner point of the second bounding box is called a second boundary corner point, including the upper left corner point, the lower left corner point, the upper right corner point, and the lower right corner point. Among them, the second boundary corner points are divided into left corner points and right corner points in the horizontal direction. The left corner points are the upper left corner point and the lower left corner point, and the right corner points are the upper right corner point and the lower right corner point. The second boundary corner points are divided into top corner points and bottom corner points in the vertical direction. The top corner points are the upper left corner point and the upper right corner point, and the bottom corner points are the lower left corner point and the lower right corner point. The second left horizontal coordinate value is the horizontal coordinate value of the left corner point of the second bounding box, that is, the horizontal coordinate values of the upper left corner point and the lower left corner point. The second right horizontal coordinate value is the horizontal coordinate value of the right corner point of the second bounding box, that is, the horizontal coordinate values of the upper right corner point and the lower right corner point. The second top vertical coordinate value is the vertical coordinate value of the top corner point of the second bounding box, that is, the vertical coordinate values of the upper left corner point and the upper right corner point. The second bottom vertical coordinate value is the vertical coordinate value of the bottom corner point of the second bounding box, that is, the vertical coordinate values of the lower left corner point and the lower right corner point.

[0084] Traverse all bounding box groups. For the second bounding box in each bounding box group, extract the coordinate information of each second boundary corner point, including the horizontal coordinate value of the left corner point, the horizontal coordinate value of the right corner point, the vertical coordinate value of the top corner point, and the vertical coordinate value of the bottom corner point. Exemplarily, rect1 and rect2 respectively represent the second bounding box and the second bounding box. The bounding box information consists of four values, which are the upper left coordinates (x, y) of the bounding box and the width (w) and height (h) of the bounding box. That is, the bounding box information of rect2 is (x2, y2) and the width w2 and height h2. Then the second left horizontal coordinate value of rect2 is x2, the second right horizontal coordinate value of rect2 is x2, x2 + w2, the second top vertical coordinate value of rect2 is y2, and the second bottom vertical coordinate value of rect2 is y2 + h2.

[0085] Step S20223, if the first right horizontal coordinate value is greater than the second left horizontal coordinate value and the second right horizontal coordinate value is greater than the first left horizontal coordinate value, and, the first bottom vertical coordinate value is greater than the second top vertical coordinate value and the second bottom vertical coordinate value is greater than the first top vertical coordinate value, then determine that the bounding box group is an overlapping box group.

[0086] Based on the extracted corner coordinate values of the first bounding box and the second bounding box, determine whether the two bounding boxes overlap. The specific conditions are: the right horizontal coordinate value of the first bounding box is greater than the left horizontal coordinate value of the second bounding box, the right horizontal coordinate value of the second bounding box is greater than the left horizontal coordinate value of the first bounding box, the bottom vertical coordinate value of the first bounding box is greater than the top vertical coordinate value of the second bounding box, and the bottom vertical coordinate value of the second bounding box is greater than the top vertical coordinate value of the first bounding box.

[0087] It is understandable that by comparing the boundary coordinates of two bounding boxes, it is determined whether they overlap in the horizontal and vertical directions. If the two bounding boxes overlap in both the horizontal and vertical directions, it is considered that they overlap, thus achieving an accurate determination of whether the two bounding boxes intersect, providing a basis for subsequent tasks such as image complexity assessment.

[0088] Exemplarily, the code for calculating the number of overlaps can be:

[0089]

[0090] Where x1, y1, w1, and h1 are the x-coordinate, y-coordinate, width data, and height data in the first bounding box information respectively; x2, y2, w2, and h2 are the x-coordinate, y-coordinate, width data, and height data in the second bounding box information respectively. That is, when the leftmost width of x1 is smaller than the rightmost width of x2, the rightmost is larger than the leftmost of x2, the lowest is lower than the highest of y2, and the highest is higher than the lowest of y2, it is determined that the two boxes overlap.

[0091] In a feasible embodiment, the step S30: determining the image complexity of the to-be-processed image based on the object distribution characteristics includes:

[0092] Step S301, obtaining a preset weight group, and performing a weighted summation process on the bounding box area, the number of box overlaps, and the number of bounding boxes of the at least one bounding box through the preset weight group to obtain a weighted processing value;

[0093] The preset weight group is a set of preset weight values. Each weight value is used to perform a weighted summation process on different features (such as the bounding box area, the number of box overlaps, and the number of bounding boxes). These weight values can be adjusted according to actual needs to reflect the relative importance of different features in image complexity assessment. The process of multiplying the bounding box area, the number of box overlaps, and the number of bounding boxes by their corresponding weight values and then adding the products. It is understandable that through weighted summation, the contributions of different features to image complexity can be comprehensively considered, improving the accuracy of image complexity.

[0094] In the specific implementation, before performing the weighted summation process, the bounding box area, the number of box overlaps, and the number of bounding boxes can be normalized to limit the bounding box area, the number of box overlaps, and the number of bounding boxes within the same value range for subsequent calculation. The specific formula for the weighted summation process can be:

[0095] weighted_score=(all_area_scaled*w1 + overlap_scaled*w2 + count_scaled*w3), where,

[0096] all_area_scaled is the normalized value of the bounding box area, overlap_scaled is the normalized value of the box overlap count, count_scaled is the normalized value of the number of bounding boxes, and w1, w2, and w3 are the weight values corresponding to the bounding box area, box overlap count, and number of bounding boxes respectively. It should be noted that the weight coefficients (w1, w2, w3) can be dynamically set according to the application scenario. For example, in the object detection scenario, emphasis is placed on overlap and distribution density, so w2 > w1 > w3; in the image stitching scenario, emphasis is placed on the area coverage range, so w1 > w2 > w3.

[0097] Step S302: Determine the image complexity of the image to be processed based on the weighted processing value, where the weighted processing value has a positive correlation with the image complexity.

[0098] Determine the image complexity according to the weighted processing value. The larger the weighted processing value, the higher the image complexity; the smaller the weighted processing value, the lower the image complexity. To determine the image complexity according to the weighted processing value, a linear or non-linear function can be used for mapping. For example, the weighted processing value can be determined as the image complexity, or the weighted processing value can be mapped to the range [0, 1] to obtain the image complexity. Specifically, it can be set according to actual needs and is not limited here. It can be understood that in this embodiment, by establishing a positive correlation between the weighted processing value and the image complexity, the complexity of the image can be intuitively reflected, providing a basis for subsequent image processing.

[0099] Based on the first and / or second embodiments of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first and / or second embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, the preset image processing is image stitching processing; after the step S30: determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, the following steps are further included:

[0100] Step S40: If the image complexity is greater than the preset complexity threshold, use a preset high-precision image stitching algorithm to perform image stitching on the image to be processed;

[0101] Image complexity can characterize the complexity reflected by comprehensive factors such as the number of objects contained in the image, the spatial relationships between objects, the texture details of objects, the richness of colors, and the complexity of the image background. The higher the image complexity, the more information and the more difficult it is to process in the image. A preset complexity threshold is used as the boundary for judging the high or low image complexity. When the image complexity is greater than this threshold, it indicates that the image is relatively complex and a high-precision image stitching algorithm needs to be adopted. The high-precision image stitching algorithm has higher precision and complexity in feature matching, image fusion, etc. It has a large feature matching density and a high complexity of image fusion.

[0102] Specifically, first, feature points are extracted from the image to be processed based on a feature extraction algorithm; a high-precision feature matching strategy is adopted to match the feature points in different images, such as two-way matching verification, random sample consensus algorithm, etc., to ensure the accuracy of the matching; a transformation matrix between the images is calculated according to the matched feature points, such as an affine transformation matrix or a perspective transformation matrix; through this transformation matrix, different images are aligned to the same coordinate system. For example, the least squares method is used to solve the transformation matrix to minimize the error between the matching points; an image fusion algorithm, such as a multi-resolution fusion algorithm, is used to fuse the aligned images. The multi-resolution fusion algorithm can decompose the image into sub-images of different scales, then perform fusion at each scale, and finally reconstruct the fused sub-images into the final stitched image. This can effectively eliminate obvious traces at the stitching location and make the stitched image have a natural transition.

[0103] It can be understood that in the stitching scenario of high-complexity images, the high-precision image stitching algorithm can provide more accurate feature matching and more natural image fusion effects, thus obtaining high-quality stitched images. For application fields that require accurate image information, such as geographical mapping and medical diagnosis, etc., it can improve work efficiency and the accuracy of decision-making.

[0104] Step S50, if the image complexity is less than or equal to the preset complexity threshold, then a preset low-precision image stitching algorithm is used to perform image stitching on the image to be processed; wherein, the feature matching density in the high-precision image stitching algorithm is greater than that in the low-precision image stitching algorithm, and / or, the complexity of image fusion in the high-precision image stitching algorithm is higher than that in the low-precision image stitching algorithm, and / or, the computational resource consumption of the high-precision image stitching algorithm is greater than that of the low-precision image stitching algorithm.

[0105] The low-precision image stitching algorithm has lower precision and complexity in feature matching, image fusion, etc. Its feature matching density is small, the requirements for feature extraction and matching of the image are relatively loose, the complexity of image fusion is low, and the computational resource consumption is also relatively small.

[0106] Specifically, a relative feature extraction algorithm, such as the ORB (Oriented FAST and Rotated BRIEF) algorithm, is used for feature extraction, and then a matching strategy, such as brute-force matching or a matching method based on the Hamming distance, is adopted to match the feature points in different images; according to the matched feature points, a transformation matrix between the images is calculated, and usually a relatively simple transformation model, such as a translation transformation or an affine transformation, is adopted; an image fusion algorithm is used for image stitching, such as a linear fusion algorithm, and linear weighted averaging is performed on the pixels at the stitching position to achieve the transition of the image.

[0107] It can be understood that in some real-time video surveillance or live broadcast scenarios, it is necessary to quickly stitch the video frames of multiple cameras to provide a panoramic view. Due to the high real-time requirement and relatively low requirement for stitching accuracy, a low-precision image stitching algorithm can complete the stitching task in a short time and meet the real-time requirement. In the stitching scenario of low-complexity images, the low-precision image stitching algorithm has the advantages of fast calculation speed and low resource consumption, and can quickly complete the stitching task on the premise of ensuring a certain stitching effect, meeting the application scenarios with high real-time requirements.

[0108] In a feasible implementation manner, the preset image processing is object detection; after the step S30: determining the image complexity of the to-be-processed image corresponding to the object distribution feature based on the object distribution feature, the following steps are further included:

[0109] Step S60, if the image complexity is greater than a preset complexity threshold, then a preset high-precision object detection algorithm is used to perform object detection on the to-be-processed image;

[0110] The high-precision object detection algorithm has a large density of sliding windows and can scan the image area more carefully; the number of scales of the detection windows is large and can cover more targets of different sizes; the computational complexity is high, and usually more computing resources and time are required to complete the detection task, but it can provide more accurate detection results.

[0111] A powerful feature extractor is used for feature extraction, such as a pre-trained model based on a deep convolutional neural network; a high-density sliding window is used to scan the feature map, and a smaller step size of the sliding window can more comprehensively cover the image area to ensure that the positions where targets may exist are not missed. At the same time, a relatively large number of detection windows of different scales are set to adapt to targets of different sizes. For example, for targets with large size differences, windows of different sizes are set for detection; for the area corresponding to each sliding window, a classifier and a regressor are used for object classification and position localization. Among them, the classifier determines whether the area contains a target and the category of the target, and the regressor predicts the specific position of the target in the image; the non-maximum suppression algorithm is used to remove the overlapping detection frames to obtain the object detection result.

[0112] It can be understood that in the target detection scenario of high-complexity images, a high-precision target detection algorithm can fully exert its advantages. Through a high-density sliding window, multi-scale detection windows, and a complex calculation process, it can accurately identify and locate targets.

[0113] Step S70: If the image complexity is less than or equal to the preset complexity threshold, a preset low-precision target detection algorithm is used to perform target detection on the image to be processed; wherein, the density of the sliding window in the high-precision target detection algorithm is greater than that in the low-precision target detection algorithm, and / or the number of scales of the detection window in the high-precision target detection algorithm is greater than that in the low-precision target detection algorithm, and / or the computational complexity of the high-precision target detection algorithm is greater than that in the low-precision target detection algorithm.

[0114] The low-precision target detection algorithm has a relatively low density of sliding windows, and the degree of detail in scanning the image area is relatively poor; the number of scales of the detection window is small, and the adaptability to targets of different sizes is weak; the computational complexity is low, the required computing resources and time are less, the processing speed is fast, but the detection accuracy is relatively low.

[0115] Specifically, in this embodiment, relatively simple feature extraction methods are used, such as HOG (Histogram of Oriented Gradients) features or LBP (Local Binary Patterns) features. These feature extraction methods are simple and fast to calculate, but the ability to express image features is relatively weak; a low-density sliding window is used to scan the image, the step size of the sliding window is large, and the scanned area coverage is relatively rough. At the same time, a small number of scales of detection windows are set, generally only several scales set for common-sized targets; for the area corresponding to each sliding window, a classifier (such as a support vector machine) is used for target classification; a non-maximum suppression algorithm is used to remove overlapping detection frames, but due to the relatively small number and degree of overlap of the detection frames, the complexity of post-processing is also low.

[0116] It can be understood that in the target detection scenario of low-complexity images, the advantage of the low-precision target detection algorithm lies in its fast processing speed and low computing resource consumption, which can complete the target detection task in a short time and meet the application scenarios with high real-time requirements.

[0117] In a feasible embodiment, the preset image processing is the image processing for robot navigation and obstacle avoidance; after the step S30: the step of determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, it further includes:

[0118] Step S80, if the complexity of the image is greater than a preset complexity threshold, then use a preset high-precision navigation algorithm to perform environmental perception and path planning on the image to be processed;

[0119] The high-precision navigation algorithm aims to perform environmental perception and path planning more precisely. The high-precision navigation algorithm has a high sensor sampling frequency and can obtain environmental information more frequently; it has a large number of modalities for sensor data fusion and comprehensively uses data from multiple sensors to improve the accuracy of environmental perception; there are many path planning constraints, considering more factors to generate a better path.

[0120] Use multiple sensors (such as lidar, cameras, depth sensors, etc.) to collect data at a high sampling frequency. For example, lidar can scan the surrounding environment dozens or even hundreds of times per second to obtain accurate distance information, and cameras can capture high-resolution images to provide rich visual information; fuse the data collected by multiple sensors. For example, combine the distance information of lidar with the visual information of cameras, and through algorithms such as Kalman filtering and extended Kalman filtering, obtain a more accurate environmental model. Specifically, data from inertial measurement units can also be fused to improve the robot's perception of its own pose and motion state; use the fused data to identify and model obstacles in the environment. Specifically, deep learning algorithms such as convolutional neural networks can be used to process camera images to identify different types of obstacles; combine lidar data to determine information such as the position, size, and shape of obstacles, and build a three-dimensional model of the environment; perform path planning based on more path planning constraints, such as the dynamic constraints of the robot, the dynamic changes of obstacles, and the accessibility of the environment. For moving obstacles in the environment, it is necessary to predict their motion trajectories and include them in the consideration of path planning; use complex path search algorithms such as the A* algorithm and the Dijkstra algorithm to search for the optimal path in the environmental model; optimize the searched path to make it smoother and more efficient.

[0121] In a high-complexity environment, the high-precision navigation algorithm can provide more accurate environmental perception and better path planning. Through high-frequency sensor sampling and multi-modal data fusion, the robot can discover obstacles and environmental changes more timely, consider more path planning constraints, and can generate safer and more efficient paths, improving the adaptability and reliability of the robot.

[0122] Step S90: If the complexity of the image is less than or equal to the complexity threshold, a preset low-precision navigation algorithm is used to perform environmental perception and path planning on the image to be processed; wherein, the sampling frequency of the sensors in the high-precision navigation algorithm is higher than that in the low-precision navigation algorithm, and / or the number of sensor data fusion modalities in the high-precision navigation algorithm is more than that in the low-precision navigation algorithm, and / or the path planning constraint conditions in the high-precision navigation algorithm are more than those in the low-precision navigation algorithm.

[0123] The low-precision navigation algorithm has the advantages of small computational complexity and fast processing speed. The sensor sampling frequency is relatively low, reducing the burden of data acquisition; the number of sensor data fusion modalities is small, perhaps only using data from some sensors; the path planning constraint conditions are few, and the process of generating a path is relatively simple.

[0124] Use fewer sensors and collect data at a relatively low sampling frequency; perform relative recognition and modeling of obstacles in the environment. For example, traditional image processing algorithms such as edge detection and threshold segmentation can be used to process camera images to identify the approximate positions of obstacles; determine the constraint conditions. The low-precision navigation algorithm considers fewer path planning constraint conditions and mainly focuses on the static positions of obstacles and the basic motion capabilities of the robot. For example, only basic constraints such as the maximum turning radius and minimum safety distance of the robot are considered; use a path search algorithm such as the greedy algorithm to find a path; during the movement of the robot, adjust the path according to real-time environmental changes. For example, if a suddenly emerging obstacle is encountered, the robot can change direction to avoid the obstacle.

[0125] In a low-complexity environment, the advantage of the low-precision navigation algorithm lies in its simplicity and efficiency. By reducing the sensor sampling frequency, reducing the data fusion modalities, and simplifying the path planning constraint conditions, the computational complexity and processing time of the algorithm are reduced, and the response speed of the robot is improved.

[0126] In a feasible implementation, the preset image processing is video surveillance and analysis processing; after the step S30: the step of determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, the following is further included:

[0127] Step A10: If the image complexity is greater than the preset complexity threshold, a preset high-precision surveillance algorithm is used to perform target detection and tracking on the image to be processed.

[0128] The characteristics of the high-precision monitoring algorithm are high detection frame rate, which can process more video frames per unit time and capture the dynamic changes of the target in a timely manner; the associated feature dimensions for tracking the target are numerous, integrating more feature information to accurately associate and track the target; and a large amount of computing resource allocation, investing more computing resources to improve the performance of the algorithm.

[0129] Process the video frames at a relatively high frame rate. For example, use object detection models based on deep learning, such as Faster R-CNN, YOLO series, etc., which can process dozens or even hundreds of images per second. In each frame, the model will comprehensively scan the image, identify the possible targets, and give the category and bounding box information of the targets; multi-feature fusion detection, by comprehensively using multiple features for object detection. For example, combine the optical flow method to extract the motion information of the target and fuse it with the appearance features extracted by the deep learning model to improve the accuracy of object detection; during the tracking process, use features in multiple dimensions to associate the targets in different frames. The feature dimensions can include the color histogram of the target, HOG features, deep learning features, etc. By calculating the similarity between different features, determine the correspondence of the targets in different frames. For example, use the Hungarian algorithm or the Kalman filter algorithm to match and track the targets according to the feature similarity and the motion trajectory of the targets; when the target is occluded, use the historical motion information and feature information of the target for prediction, and re-associate the target after the occlusion is removed; when the target disappears and reappears briefly, set a certain time window and feature matching threshold to determine whether it is the same target.

[0130] In high-complexity video surveillance scenarios, the high-precision monitoring algorithm can accurately detect and track targets by virtue of its high detection frame rate, multi-dimensional feature association, and large amount of computing resource investment. Even in the case of dense targets, frequent occlusions, and complex backgrounds, it can provide reliable monitoring information, helping to detect abnormal events in a timely manner and ensuring public safety and improving management efficiency.

[0131] Step A20, if the image complexity is less than or equal to the complexity threshold, then use a preset low-precision monitoring algorithm to perform object detection and tracking on the image to be processed; where the detection frame rate in the high-precision monitoring algorithm is higher than that in the low-precision monitoring algorithm, and / or the associated feature dimensions for tracking the target in the high-precision monitoring algorithm are more than those in the low-precision monitoring algorithm, and / or the computing resource allocation amount in the high-precision monitoring algorithm is greater than that in the low-precision monitoring algorithm.

[0132] The detection frame rate of the low-precision monitoring algorithm is low, and the number of video frames processed per unit time is small; the associated feature dimensions for tracking the target are few, mainly relying on a small number of key features for target association; the computing resource allocation amount is small, and the requirements for hardware resources are low.

[0133] Process video frames at a lower frame rate. For example, process a few frames of images per second to reduce the computational load. Some lightweight object detection algorithms can be used, such as cascade classifiers based on Haar features, machine learning classifiers, etc.; perform object detection based on single or a small number of features. For example, only use color features or shape features to identify objects, and by setting a certain feature threshold, determine whether there are objects in the image; during the tracking process, use fewer feature dimensions to associate objects in different frames. The feature dimensions can include the position and appearance features of the object, and by calculating the distance of the object position and the similarity of the appearance features, determine whether the objects in different frames are the same object.

[0134] In low-complexity video surveillance scenarios, the advantage of low-precision surveillance algorithms lies in their simplicity and efficiency. By reducing the detection frame rate, reducing the feature dimensions, and allocating computing resources, the computational load of the algorithm is reduced, the processing speed is increased, and at the same time, the requirements for hardware devices are reduced, enabling the surveillance system to achieve basic object detection and tracking functions at low cost, and being suitable for application scenarios where the requirements for surveillance accuracy are not high but are sensitive to cost and resource consumption.

[0135] In a feasible implementation manner, the preset image processing is image generation and editing processing; after the step S30: determining the image complexity of the to-be-processed image corresponding to the object distribution feature based on the object distribution feature, the following is further included:

[0136] Step A30, if the image complexity is greater than a preset complexity threshold, then use a preset high-precision generation algorithm to perform image generation or editing on the to-be-processed image;

[0137] The generation model of the high-precision generation algorithm has a large network depth and can learn more complex features and patterns; the discriminator has a large number of stages and can more finely judge the authenticity of the generated image; there are many diversity constraint conditions for scene elements, ensuring that the generated image can cover a variety of scene elements and be reasonably combined.

[0138] The high-precision generation algorithm can use a deep neural network as the generation model, such as a deep extended version of the Deep Convolutional Generative Adversarial Network (DCGAN) or a generation model based on the Transformer architecture. Increasing the network depth means having more convolutional layers or Transformer blocks, which can extract and learn more abstract and complex features of the image. For example, when generating an image containing complex buildings and natural landscapes, a deep network can learn features such as the structural details of the buildings and the texture changes of the natural landscapes. When editing an image, the model can precisely modify the image according to the input editing instructions by using the learned features; judge the generated image at multiple levels and scales, and each stage can focus on different features, such as local details, overall structure, etc. For example, the discriminator in the first stage can check whether the local texture of the image is real, and the discriminator in the second stage can judge the overall layout and scene rationality of the image. Through multiple stages of discrimination, the generation model can be more accurately guided to generate high-quality images; during the generation or editing process, set more constraints on the diversity of scene elements. For example, it is stipulated that the generated image contains a certain number and types of objects, and the distribution of these objects should conform to a certain rule. When editing an image, it is required that the newly added elements be coordinated with the elements in the original image in terms of style and layout; introduce reinforcement learning into the high-precision generation algorithm. By designing a suitable reward function, the generation model can continuously optimize its generation strategy in the interaction with the environment (discriminator). For example, the reward function can consider factors such as the diversity of the generated image and the degree of compliance with the editing instructions, prompting the generation model to generate higher-quality images; multi-modal fusion, that is, fusing information of images and other modalities, such as text descriptions, audio information, etc. For example, generate or edit an image according to the text description to make the image more in line with the specific needs of the user. During the editing process, combining audio information can add dynamic effects or a specific atmosphere to the image.

[0139] In the scenarios of generating and editing high-complexity images, the high-precision generation algorithm, relying on its deep generation model, multi-stage discriminator, and diverse scene element constraint conditions, can generate or edit high-quality, detail-rich images that meet the requirements of complex scenarios, improving the quality and visual effects of the works and meeting the user's needs for high-quality images.

[0140] Step A40: If the image complexity is less than or equal to the complexity threshold, use a preset low-precision generation algorithm to perform image generation or editing on the image to be processed; wherein, the network depth of the generation model in the high-precision generation algorithm is greater than that in the low-precision generation algorithm, and / or the number of stages of the discriminator in the high-precision generation algorithm is more than that in the low-precision generation algorithm, and / or the diversity constraint conditions of the scene elements in the high-precision generation algorithm are more than those in the low-precision generation algorithm.

[0141] The network depth of the generation model of the low-precision generation algorithm is relatively shallow, and its learning ability is relatively weak; the discriminator has a small number of stages, and the judgment of image authenticity is relatively rough; there are few diversity constraints on scene elements, and the generated images are relatively poor in terms of the richness of scene elements and the rationality of layout.

[0142] In this embodiment, a relatively shallow neural network is used as the generation model, such as a multi-layer perceptron (MLP) or a shallow convolutional neural network. This model has a small number of parameters and a small amount of computation, and can generate images quickly; judgment is based on a small number of discriminators. For example, it focuses on whether the overall color distribution and basic shape of the image are reasonable, without deeply examining local details; fewer diversity constraints on scene elements are set. For example, the generated image contains a small number of basic elements, and there are no strict requirements for the distribution and combination of elements. When editing the image, only some modifications are made, such as adjusting the color, brightness, etc. It can be understood that in the generation and editing scenarios of low-complexity images, the low-precision generation algorithm has the advantages of fast calculation speed and low resource consumption, and can complete the image generation or editing task in a short time, meeting the user's demand for quick feedback.

[0143] Exemplarily, to help understand the implementation process of the image processing method obtained by combining the above embodiments with this embodiment, please refer to Figure 3 , Figure 3 A brief flow schematic diagram of an image processing method is provided. Specifically:

[0144] S1. Use the SAM model to segment the image to be processed, generating multiple bounding boxes (that is, perform image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box). The segmentation result of the SAM model can be as Figure 4As shown. S2. Calculate the total area of the regions corresponding to all the bounding boxes to obtain the bounding box area, calculate the number of times the bounding boxes overlap with each other, and count the number of bounding boxes (that is, determine the number of bounding boxes of the at least one bounding box, and determine the bounding box area of the at least one bounding box based on the bounding box information; determine the number of times the bounding boxes overlap based on the bounding box information, where the number of times the bounding boxes overlap is used to represent the number of times any two of the at least one bounding box overlap; determine the number of bounding boxes, the bounding box area, and the number of times the bounding boxes overlap as the object distribution characteristics of the image to be processed). S3. Perform a weighted summation process on the bounding box area, the number of times the bounding boxes overlap, and the number of bounding boxes through a preset weight group to obtain a weighted processing value, and use the weighted processing value as the image complexity (that is, determine the image complexity of the image to be processed based on the object distribution characteristics). S4. Perform image processing on the image to be processed based on the image complexity (that is, perform image processing on the image to be processed based on the image complexity).

[0145] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image processing method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0146] This application also provides an image processing device. Please refer to Figure 5 , the image processing device includes:

[0147] An image segmentation module 10, configured to perform image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box, where the bounding box information is used to represent the position and size of the object corresponding to the at least one bounding box in the image to be processed;

[0148] A feature determination module 20, configured to determine the object distribution characteristics of the image to be processed based on the bounding box information of the at least one bounding box, where the object distribution characteristics are used to represent the spatial distribution state of the objects in the image to be processed in the image to be processed;

[0149] A complexity determination module 30, configured to determine the image complexity of the image to be processed corresponding to the object distribution characteristics based on the object distribution characteristics for performing preset image processing on the image to be processed.

[0150] Optionally, the feature determination module 20 is further configured to:

[0151] Determine the number of bounding boxes of the at least one bounding box, and determine the bounding box area of the at least one bounding box based on the bounding box information;

[0152] Determine the number of box overlaps of the at least one bounding box based on the bounding box information, where the number of box overlaps is used to characterize the number of times of overlap between any two of the at least one bounding box;

[0153] Determine the number of bounding boxes, the area of the bounding boxes, and the number of box overlaps as the object distribution characteristics of the image to be processed.

[0154] Optionally, the number of the at least one bounding box is greater than one; the feature determination module 20 is further configured to:

[0155] Divide each bounding box into multiple groups of bounding box groups, where each group of bounding box groups includes two of the bounding boxes and each group of bounding box groups does not repeat;

[0156] Determine overlapping box groups from multiple groups of the bounding box groups based on the bounding box information, and determine the number of overlapping box groups in the multiple groups of the bounding box groups as the number of box overlaps of each of the bounding boxes, where the overlapping box group is a group of bounding box groups in which there is an overlapping area between two of the bounding boxes in the group.

[0157] Optionally, the group of bounding box groups includes a first bounding box and a second bounding box; optionally, the feature determination module 20 is further configured to:

[0158] Traverse each group of the bounding box groups, and based on the first bounding box information of the first bounding box, determine the first left abscissa value of the left corner point among each first boundary corner point, and determine the first right abscissa value of the right corner point among each of the first boundary corner points, and determine the first top ordinate value of the top corner point among each of the first boundary corner points, and determine the first bottom ordinate value of the bottom corner point among each of the first boundary corner points, where the first boundary corner point is each corner point of the first bounding box;

[0159] Based on the second bounding box information of the second bounding box, determine the second left abscissa value of the left corner point among each second boundary corner point, and determine the second right abscissa value of the right corner point among each of the second boundary corner points, and determine the second top ordinate value of the top corner point among each of the second boundary corner points, and determine the second bottom ordinate value of the bottom corner point among each of the second boundary corner points, where the second boundary corner point is each corner point of the second bounding box;

[0160] If the first right abscissa value is greater than the second left abscissa value and the second right abscissa value is greater than the first left abscissa value, and, the first bottom ordinate value is greater than the second top ordinate value and the second bottom ordinate value is greater than the first top ordinate value, then determine that the group of bounding box groups is an overlapping box group.

[0161] Optionally, the complexity determination module 30 is further configured to:

[0162] Obtain a preset weight combination, and perform a weighted summation process on the number of bounding boxes, the area of the bounding boxes, and the number of box overlaps through the preset weight combination to obtain a weighted processing value;

[0163] Determine the image complexity of the image to be processed based on the weighted processing value, where the weighted processing value has a positive correlation with the image complexity.

[0164] Optionally, the image segmentation module 10 is further configured to:

[0165] Input the image to be processed into a preset image segmentation model to obtain at least one bounding box and the bounding box information of the at least one bounding box.

[0166] Optionally, the image segmentation model is a SAM model, and the SAM model includes an image encoder, a prompt encoder, and a decoder; the image segmentation module 10 is further configured to:

[0167] Input the image to be processed into the image encoder to obtain a first output data;

[0168] Receive user prompt information, and input the user prompt information into the prompt encoder to obtain a second output data;

[0169] Input the first output data and the second output data into the decoder to obtain a segmentation mask;

[0170] Traverse each of the segmentation masks, determine the contour coordinates of the segmentation region corresponding to the segmentation mask in the image to be processed, generate a bounding box based on the contour coordinates, and determine the bounding box information of the bounding box based on the contour coordinates.

[0171] Optionally, the preset image processing is image stitching processing; the device further includes an image processing module, which is used for:

[0172] If the image complexity is greater than a preset complexity threshold, perform image stitching on the image to be processed using a preset high-precision image stitching algorithm;

[0173] If the image complexity is less than or equal to the preset complexity threshold, perform image stitching on the image to be processed using a preset low-precision image stitching algorithm; wherein, the density of feature matching in the high-precision image stitching algorithm is greater than that in the low-precision image stitching algorithm, and / or, the complexity of image fusion in the high-precision image stitching algorithm is higher than that in the low-precision image stitching algorithm, and / or, the computational resource consumption of the high-precision image stitching algorithm is greater than that in the low-precision image stitching algorithm.

[0174] Optionally, the preset image processing is object detection; the image processing module of the device is further configured to:

[0175] If the complexity of the image is greater than a preset complexity threshold, use a preset high-precision object detection algorithm to perform object detection on the image to be processed;

[0176] If the complexity of the image is less than or equal to the preset complexity threshold, use a preset low-precision object detection algorithm to perform object detection on the image to be processed; wherein, the density of the sliding window in the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm, and / or the number of scales of the detection window in the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm, and / or the computational complexity of the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm.

[0177] Optionally, the preset image processing is image processing for robot navigation and obstacle avoidance; the image processing module of the device is further configured to:

[0178] If the complexity of the image is greater than a preset complexity threshold, use a preset high-precision navigation algorithm to perform environmental perception and path planning on the image to be processed;

[0179] If the complexity of the image is less than or equal to the complexity threshold, use a preset low-precision navigation algorithm to perform environmental perception and path planning on the image to be processed; wherein, the sampling frequency of the sensor in the high-precision navigation algorithm is higher than that in the low-precision navigation algorithm, and / or the number of modalities of sensor data fusion in the high-precision navigation algorithm is more than that in the low-precision navigation algorithm, and / or the path planning constraint conditions in the high-precision navigation algorithm are more than those in the low-precision navigation algorithm.

[0180] Optionally, the preset image processing is video surveillance and analysis processing; the image processing module of the device is further configured to:

[0181] If the complexity of the image is greater than a preset complexity threshold, use a preset high-precision surveillance algorithm to perform object detection and tracking on the image to be processed;

[0182] If the complexity of the image is less than or equal to the complexity threshold, use a preset low-precision surveillance algorithm to perform object detection and tracking on the image to be processed; wherein, the detection frame rate in the high-precision surveillance algorithm is higher than that in the low-precision surveillance algorithm, and / or the number of associated feature dimensions of the tracking target in the high-precision surveillance algorithm is more than that in the low-precision surveillance algorithm, and / or the amount of computational resources allocated in the high-precision surveillance algorithm is greater than that in the low-precision surveillance algorithm.

[0183] Optionally, the preset image processing is image generation and editing processing; the image processing module is further configured to:

[0184] If the complexity of the image is greater than a preset complexity threshold, a preset high-precision generation algorithm is used to perform image generation or editing on the image to be processed;

[0185] If the complexity of the image is less than or equal to the complexity threshold, a preset low-precision generation algorithm is used to perform image generation or editing on the image to be processed; wherein, the network depth of the generation model in the high-precision generation algorithm is greater than that in the low-precision generation algorithm, and / or, the number of stages of the discriminator in the high-precision generation algorithm is more than that in the low-precision generation algorithm, and / or, the diversity constraint conditions of the scene elements in the high-precision generation algorithm are more than those in the low-precision generation algorithm.

[0186] The image processing device provided by the present application adopts the image processing method in the above embodiment, and can solve the technical problem that the large amount of calculation in the image complexity calculation method leads to waste of computing resources. Compared with the prior art, the beneficial effects of the image processing device provided by the present application are the same as those of the image processing method provided by the above embodiment, and other technical features in the image processing device are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.

[0187] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method in the first embodiment above.

[0188] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.

[0189] As Figure 6As shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0190] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0191] The electronic device provided in the present application adopts the image processing method in the above embodiment, and can solve the technical problem that the large amount of calculation in the image complexity calculation method leads to waste of calculation resources. Compared with the prior art, the beneficial effects of the electronic device provided in the present application are the same as those of the image processing method provided in the above embodiment, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0192] It should be understood that each part disclosed in the present application may be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0193] As described above, this is only the specific implementation of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

[0194] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image processing method in the above embodiments.

[0195] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0196] The above computer-readable storage medium can be included in an electronic device; or it can exist separately without being assembled into the electronic device.

[0197] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the image processing method in the above embodiments.

[0198] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0199] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0200] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0201] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned image processing method, and can solve the technical problem of large computational amount in the image complexity calculation method, resulting in waste of computing resources. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the image processing method provided in the above embodiments, and will not be elaborated here.

[0202] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the image processing method as described above.

[0203] The computer program product provided by the present application can solve the technical problem that the large amount of calculation in the image complexity calculation method leads to waste of computing resources. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the image processing method provided by the above embodiments, and will not be elaborated here.

[0204] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. An image processing method, characterized in that, The described image processing method includes: Performing image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box; wherein, the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed; Determining the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box, wherein the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed; Determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature for performing preset image processing on the image to be processed.

2. The image processing method according to claim 1, wherein The step of determining the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box includes: Determining the number of bounding boxes of the at least one bounding box and determining the bounding box area of the at least one bounding box based on the bounding box information; Determining the number of box overlaps of the at least one bounding box based on the bounding box information, wherein the number of box overlaps is used to characterize the number of times of overlap between any two bounding boxes in the at least one bounding box; Determining the number of bounding boxes, the bounding box area, and the number of box overlaps as the object distribution feature of the image to be processed.

3. The image processing method according to claim 2, wherein The number of the at least one bounding box is greater than one; The step of determining the number of box overlaps of the at least one bounding box based on the bounding box information includes: Dividing each bounding box into multiple groups of bounding box groups, wherein each bounding box group includes two of the bounding boxes and each group of the bounding box groups is non-repetitive; Determining the overlapping box groups from the multiple groups of bounding box groups based on the bounding box information and determining the number of overlapping box groups in the multiple groups of bounding box groups as the number of box overlaps of each of the bounding boxes, wherein the overlapping box group is a bounding box group in which there is an overlapping area between the two bounding boxes within the group.

4. The image processing method according to claim 3, characterized in that The bounding box group includes a first bounding box and a second bounding box; The step of determining the overlapping box groups from the multiple groups of bounding box groups based on the bounding box information includes: Traversing each group of the bounding box groups, determining the first left abscissa value of the left corner point among each of the first boundary corner points based on the first bounding box information of the first bounding box, and determining the first right abscissa value of the right corner point among each of the first boundary corner points, and determining the first top ordinate value of the top corner point among each of the first boundary corner points, and determining the first bottom ordinate value of the bottom corner point among each of the first boundary corner points, wherein the first boundary corner point is each corner point of the first bounding box; Determining the second left abscissa value of the left corner point among each of the second boundary corner points based on the second bounding box information of the second bounding box, and determining the second right abscissa value of the right corner point among each of the second boundary corner points, and determining the second top ordinate value of the top corner point among each of the second boundary corner points, and determining the second bottom ordinate value of the bottom corner point among each of the second boundary corner points, wherein the second boundary corner point is each corner point of the second bounding box; If the first right abscissa value is greater than the second left abscissa value, the second right abscissa value is greater than the first left abscissa value, the first bottom ordinate value is greater than the second top ordinate value, and the second bottom ordinate value is greater than the first top ordinate value, then determine that the set of bounding boxes is an overlapping box set.

5. The image processing method according to claim 2, wherein The step of determining the image complexity of the image to be processed based on the object distribution characteristics includes: Obtain a preset weight set, and perform a weighted summation process on the number of bounding boxes, the area of the bounding boxes, and the number of box overlaps through the preset weight set to obtain a weighted processing value; Determine the image complexity of the image to be processed based on the weighted processing value, where the weighted processing value has a positive correlation with the image complexity.

6. The image processing method according to claim 1, wherein The step of performing image segmentation processing on the image to be processed to obtain at least one bounding box and the bounding box information of the at least one bounding box includes: Input the image to be processed into a preset image segmentation model to obtain at least one bounding box and the bounding box information of the at least one bounding box.

7. The image processing method according to claim 6, wherein The image segmentation model is a SAM model, and the SAM model includes an image encoder, a prompt encoder, and a decoder; The step of inputting the image to be processed into a preset image segmentation model to obtain at least one bounding box and the bounding box information of the at least one bounding box includes: Input the image to be processed into the image encoder to obtain first output data; Receive user prompt information, and input the user prompt information into the prompt encoder to obtain second output data; Input the first output data and the second output data into the decoder to obtain a segmentation mask; Traverse each segmentation mask, determine the contour coordinates of the segmentation region corresponding to the segmentation mask in the image to be processed, generate a bounding box based on the contour coordinates, and determine the bounding box information of the bounding box based on the contour coordinates.

8. The image processing method according to any one of claims 1 to 7, characterized in that The preset image processing is image stitching processing; After the step of determining the image complexity of the image to be processed corresponding to the object distribution characteristics based on the object distribution characteristics, it further includes: If the image complexity is greater than a preset complexity threshold, then perform image stitching on the image to be processed using a preset high-precision image stitching algorithm; If the image complexity is less than or equal to the preset complexity threshold, then perform image stitching on the image to be processed using a preset low-precision image stitching algorithm; where the density of feature matching in the high-precision image stitching algorithm is greater than that in the low-precision image stitching algorithm, and / or the complexity of image fusion in the high-precision image stitching algorithm is higher than that in the low-precision image stitching algorithm, and / or the computational resource consumption of the high-precision image stitching algorithm is greater than that in the low-precision image stitching algorithm.

9. The image processing method according to any one of claims 1 to 7, characterized in that, The preset image processing is object detection; After the step of determining the image complexity of the image to be processed corresponding to the object distribution characteristics based on the object distribution characteristics, it further includes: If the image complexity is greater than a preset complexity threshold, then perform object detection on the image to be processed using a preset high-precision object detection algorithm; If the complexity of the image is less than or equal to the preset complexity threshold, a preset low-precision object detection algorithm is used to perform object detection on the image to be processed; wherein, the density of the sliding window in the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm, and / or the number of scales of the detection window in the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm, and / or the computational complexity of the high-precision object detection algorithm is greater than that in the low-precision object detection algorithm.

10. The image processing method according to any one of claims 1 to 7, characterized in that, The preset image processing is for image processing of robot navigation and obstacle avoidance; After the step of determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, the following is further included: If the image complexity is greater than the preset complexity threshold, a preset high-precision navigation algorithm is used to perform environmental perception and path planning on the image to be processed; If the image complexity is less than or equal to the complexity threshold, a preset low-precision navigation algorithm is used to perform environmental perception and path planning on the image to be processed; wherein, the sampling frequency of the sensor in the high-precision navigation algorithm is higher than that in the low-precision navigation algorithm, and / or the number of modalities of sensor data fusion in the high-precision navigation algorithm is more than that in the low-precision navigation algorithm, and / or the path planning constraint conditions in the high-precision navigation algorithm are more than those in the low-precision navigation algorithm.

11. The image processing method according to any one of claims 1 to 7, characterized in that, The preset image processing is for video surveillance and analysis processing; After the step of determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, the following is further included: If the image complexity is greater than the preset complexity threshold, a preset high-precision surveillance algorithm is used to perform object detection and tracking on the image to be processed; If the image complexity is less than or equal to the complexity threshold, a preset low-precision surveillance algorithm is used to perform object detection and tracking on the image to be processed; wherein, the detection frame rate in the high-precision surveillance algorithm is higher than that in the low-precision surveillance algorithm, and / or the number of associated feature dimensions of the tracking target in the high-precision surveillance algorithm is more than that in the low-precision surveillance algorithm, and / or the amount of computational resources allocated in the high-precision surveillance algorithm is greater than that in the low-precision surveillance algorithm.

12. The image processing method according to any one of claims 1 to 7, characterized in that, The preset image processing is for image generation and editing processing; After the step of determining the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, the following is further included: If the image complexity is greater than the preset complexity threshold, a preset high-precision generation algorithm is used to perform image generation or editing on the image to be processed; If the image complexity is less than or equal to the complexity threshold, a preset low-precision generation algorithm is used to perform image generation or editing on the image to be processed; wherein, the network depth of the generation model in the high-precision generation algorithm is greater than that in the low-precision generation algorithm, and / or the number of stages of the discriminator in the high-precision generation algorithm is more than that in the low-precision generation algorithm, and / or the diversity constraint conditions of the scene elements in the high-precision generation algorithm are more than those in the low-precision generation algorithm.

13. An image processing apparatus, characterized in that, The image processing device includes: An image segmentation module, configured to perform image segmentation processing on the image to be processed, to obtain at least one bounding box and the bounding box information of the at least one bounding box, wherein the bounding box information is used to characterize the position and size of the object corresponding to the at least one bounding box in the image to be processed; A feature determination module, configured to determine the object distribution feature of the image to be processed based on the bounding box information of the at least one bounding box, wherein the object distribution feature is used to characterize the spatial distribution state of the objects in the image to be processed in the image to be processed; A complexity determination module, configured to determine the image complexity of the image to be processed corresponding to the object distribution feature based on the object distribution feature, for performing preset image processing on the image to be processed.

14. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the image processing method according to any one of claims 1 to 12.

15. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the image processing method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image processing method according to any one of claims 1 to 12.