An image processing method, apparatus, electronic device, and storage medium
By segmenting and recognizing panoramic images, and using a preset segmentation model to determine the position of target feature elements and perform distortion correction, the problem of slow panoramic image processing speed in existing technologies is solved, and fast and effective image processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ARASHI VISION INC
- Filing Date
- 2022-04-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies involve cumbersome recognition and processing when capturing panoramic images, resulting in slow processing speeds.
A pre-defined segmentation model is used to segment the panoramic image to obtain a first local image with a lower degree of distortion. The position of the target feature elements is determined by feature recognition processing, and the positional relationship of the panoramic image is determined by the pre-defined segmentation model to correct the distortion.
By quickly identifying and correcting target feature elements in panoramic images, the slow processing speed of existing technologies is improved, thus increasing the efficiency of panoramic image processing.
Smart Images

Figure CN116958164B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] A panoramic image is a wide-angle image that uses wide-angle techniques and formats such as photographs and videos to represent as much of the surrounding environment as possible. A panoramic image is an image that can wrap around the viewer horizontally from 0 to 360° and vertically from 0 to 180°. Currently, when capturing panoramic images, it is necessary to perform image recognition processing to identify the subject within the panoramic image, thereby obtaining an image with the subject as the focal point. Summary of the Invention
[0003] This application provides an image processing method, apparatus, electronic device, and storage medium that can quickly identify targets in panoramic images, improving upon the problem that existing identification processes are cumbersome and result in slow processing speeds.
[0004] On one hand, embodiments of this application provide an image processing method, including:
[0005] Acquire panoramic images;
[0006] The panoramic image is segmented using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image.
[0007] The first local image is subjected to feature recognition processing to obtain the image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements;
[0008] A target feature element is determined based on multiple feature elements, and the position of the target feature element in the first local image is determined.
[0009] Based on the preset segmentation model, the positional relationship between the first local image and the panoramic image is determined, as well as the position of the target feature element in the first local image, and the position of the target feature element in the panoramic image is determined.
[0010] Based on the position of the target feature element in the panoramic image, distortion correction processing is performed on the panoramic image to obtain a corrected panoramic image.
[0011] On the other hand, embodiments of this application also provide an image processing apparatus, including:
[0012] The acquisition unit is used to acquire panoramic images;
[0013] The segmentation processing unit is used to segment the panoramic image using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image.
[0014] The feature recognition processing unit is used to perform feature recognition processing on the first local image to obtain the image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements;
[0015] The first determining unit is configured to determine a target feature element based on the plurality of feature elements, and to determine the position of the target feature element in the first local image;
[0016] The second determining unit is used to determine the positional relationship between the first local image and the panoramic image, as well as the position of the target feature element in the first local image, and the position of the target feature element in the panoramic image, based on the preset segmentation model.
[0017] The correction unit is used to perform distortion correction processing on the panoramic image based on the position of the target feature element in the panoramic image to obtain a corrected panoramic image.
[0018] On the other hand, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute steps in any of the image processing methods provided in embodiments of this application.
[0019] On the other hand, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the image processing methods provided in embodiments of this application.
[0020] In this application, a preset segmentation model is used to segment the panoramic image, thereby obtaining a first local image with a lower degree of distortion. Then, feature recognition processing is performed on the image features of the first local image to determine the target feature elements in the image features of the first local image and the position of the target feature elements in the first local image. Based on the positional change relationship corresponding to the preset segmentation model, the position of the target feature elements in the panoramic image is determined. Based on the position of the target feature elements in the panoramic image, distortion correction processing is performed on the panoramic image to obtain a corrected panoramic image. Thus, by using a preset segmentation model to quickly process the panoramic image, the problem of large distortion in some important scenes or people in the image is improved. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of a scene illustrating the image processing method provided in an embodiment of this application;
[0023] Figure 2 This is a schematic flowchart of the image processing method provided in the embodiments of this application;
[0024] Figure 2a This is a schematic diagram of a panoramic image segmented according to a preset segmentation model, provided in an embodiment of this application.
[0025] Figure 2b This is a schematic diagram of extracting a sub-image from a first partial image according to an embodiment of this application;
[0026] Figure 3 This is a flowchart illustrating the method for obtaining image features of a first local image according to an embodiment of this application;
[0027] Figure 4 This is a flowchart illustrating the method for determining target feature elements and their positions in a first partial image, as provided in an embodiment of this application.
[0028] Figure 5 This is a flowchart illustrating the method for obtaining a corrected next frame panoramic image provided in an embodiment of this application;
[0029] Figure 6 This is a schematic diagram illustrating the application of the image processing method provided in this application embodiment in a server scenario;
[0030] Figure 7 This is a schematic diagram of the first structure of the image processing apparatus provided in the embodiments of this application;
[0031] Figure 8 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] This application provides an image processing method, apparatus, storage medium, and storage medium.
[0034] Specifically, the image processing device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet, smart Bluetooth device, laptop, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.
[0035] In some embodiments, the image processing apparatus may also be integrated into multiple electronic devices, such as multiple servers, with the image processing method of this application being implemented by multiple servers.
[0036] In some embodiments, the server may also be implemented as a terminal.
[0037] For example, refer to Figure 1 The electronic device can be a server, which integrates an image processing device. In this embodiment, the server is used to acquire a panoramic image; to segment the panoramic image using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image; to perform feature recognition processing on the first local image to obtain image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements; to extract the feature elements in the image features of the first local image to determine the target feature element and the position of the target feature element in the first local image; to determine the positional change relationship corresponding to the preset segmentation model; to determine the position of the target feature element in the panoramic image based on the positional change relationship of the target feature element in the first local image; and to perform distortion correction processing on the panoramic image based on the position of the target feature element in the panoramic image to obtain a corrected panoramic image.
[0038] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0039] In this embodiment, an image processing method is provided, such as... Figure 2 As shown, the specific process of this image processing method can be as follows:
[0040] S110, Obtain panoramic image.
[0041] A panoramic image is a wide-angle image that uses wide-angle techniques and formats such as photographs and videos to represent as much of the surrounding environment as possible. A panoramic image is an image that can wrap around the viewer horizontally from 0 to 360° and vertically from 0 to 180°. Panoramic images can be captured with a regular camera and then composited, or they can be captured directly with a panoramic camera.
[0042] The panoramic image can be obtained from the current frame of a panoramic video or from a pre-captured panoramic image.
[0043] In some embodiments, the acquired panoramic image can be a rectangular image with an aspect ratio of 2:1.
[0044] S120. The panoramic image is segmented using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image.
[0045] A preset segmentation model can refer to the process of dividing a panoramic image into regions with high and low distortion using machine learning or similar methods, in order to extract the regions with low distortion. Image distortion refers to the blurring, stretching, and deformation of figures or objects in an image during imaging, which prevents the image from accurately reflecting its true state. The degree of distortion can be judged by human experience or by comparing the coordinates of the distorted location with those under normal conditions. Region division can refer to dividing the panoramic image into several regions based on the degree of distortion. In some embodiments, the panoramic image can be divided into two parts based on the degree of distortion. For example, the region with a distortion degree lower than a preset distortion degree is the low-distortion region, i.e., the region containing the first local image; the region with a distortion degree higher than the preset distortion degree is the high-distortion region, i.e., the region containing the second local image.
[0046] In this embodiment, the preset segmentation model can be defined by dividing the panoramic image into low-distortion and high-distortion regions based on human experience, for example, such as... Figure 2a The panoramic image shown has a low-distortion region that can be the 80% of the image area located in the middle of the image along the width direction. Figure 2a The rectangular frame in the center of the panoramic image shown is located in the middle of the image and has a better imaging effect. The high distortion area of the panoramic image can be the upper and lower 10% of the image area, i.e. Figure 2a The upper and lower rectangular frames of the panoramic image shown are located at the upper and lower edges of the panoramic image, resulting in poor image quality. A preset segmentation model can be used to quickly process the panoramic image.
[0047] Segmentation processing can divide a panoramic image into several images with different degrees of distortion according to the division result of a preset segmentation model. For example, in the embodiments of this application, the preset segmentation model divides the panoramic image into a low-distortion region and a high-distortion region, and the segmentation processing divides the panoramic image into a first local image of the low-distortion region and a second local image of the high-distortion region.
[0048] The first and second local images can be obtained either through sliding window sampling or by segmenting the panoramic image. The number of first and second local images can be arbitrary; there can be no overlap between the first and second local images; and the first and second local images can be stitched together to form a panoramic image.
[0049] S130. Perform feature recognition processing on the first local image to obtain the image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements.
[0050] Feature recognition processing can refer to pixel-level image classification of a first local image, labeling and determining the object category to which each pixel in the first local image belongs, thereby obtaining images with different object category labels. Object categories can include landscapes, buildings, and people. After labeling the object categories on the first local image, the resulting images with different labels are image features. In some embodiments, image features can include color features, texture features, shape features, and spatial relationship features. In some embodiments, image features can be composed of multiple feature elements, where the feature elements constituting the image features can be pixels. A pixel refers to the basic encoding of a primary color and its grayscale. Pixel category is also known as pixel type; different images have different pixel types, but different values need to be passed to the template parameters for different pixel types. For example, pixel data types include CV_32U, CV_32S, CV_32F, CV_8U, CV_8UC3, etc.
[0051] For example, in this embodiment of the application, by feature recognition processing, the scenery, buildings and people in the first partial image can be identified, and by labeling and segmenting the scenery, buildings and people in the first partial image, the background part and the foreground part of the first partial image can be determined, and can be represented in the form of a mask image, wherein the mask image refers to presenting the background and foreground (the foreground may include buildings and people) of the first partial image through images with different gray levels.
[0052] Among them, such as Figure 3 As shown, in some embodiments, the method for obtaining image features of the first local image includes:
[0053] S131. Determine the size information of the first partial view.
[0054] The size information may include the length, width, and angular field of view of the first partial image. The field of view (FOV), also known as the field of view in optical engineering, determines the field of view of an optical instrument. In this embodiment, when the captured image is a panoramic image, the panoramic image has a horizontal field of view of 0–360° and a vertical field of view of 0–180°.
[0055] The size information of the first local image can be determined by measuring the first local image or by determining the segmentation ratio of the preset segmentation model.
[0056] S132. Based on the size information of the first partial view, determine the size, sliding step size, and sliding direction of the sliding window.
[0057] A sliding window, also known as a sliding frame window, is used to improve data accuracy by expanding the range of values at a given point to include that point within a larger region. This region is then used for calculations; the region is the window. A sliding window can frame a time series data within a specified unit length to calculate statistical indicators. It's analogous to a slider of a specified length moving on a ruler; each unit movement provides feedback on the image data within the slider.
[0058] The size of the sliding window is the range of image segmentation performed on the first partial image. The sliding window can be of any shape. In some embodiments, the sliding window can be a rectangular frame, and the size of the sliding window can include the length and width of the rectangular frame. In some embodiments, the sliding window can be a square frame, and the size of the sliding window can include the side length of the square frame.
[0059] The sliding direction of the sliding window is the direction in which the sliding window slides when the first partial view is slid, and the sliding direction of the sliding window can be arbitrary. In some embodiments, the sliding direction of the sliding window is along the length direction of the first partial view.
[0060] The sliding step size of the sliding window is the distance the sliding window slides each time it slides in the first partial view. The sliding step size of the sliding window can be set according to the size information of the first partial view. In some embodiments, when the sliding direction of the sliding window is along the length direction of the first partial view, the sliding step size of the sliding window is equal to the length of the first partial view divided by the number of times the sliding window slides.
[0061] S133. Based on the size of the sliding window, the sliding step size, and the sliding direction, the first local image is segmented to obtain multiple sub-images.
[0062] Segmentation refers to extracting the image from the location of the sliding window into a first local image.
[0063] The segmentation of the first local image can be performed by capturing a separate image, denoted as a sub-image, each time the sliding window moves along a preset direction and slides over a certain number of steps. When the sliding window slides over n steps, n sub-images are obtained. These sub-images can be set to alternate, spaced apart, or partially overlap.
[0064] For example, in this embodiment of the application, the preset direction is along the length direction of the first partial image. The sliding window slides through two sliding steps, resulting in two sub-images. The sum of the lengths of the two sub-images is greater than the length of the first partial image, i.e. Figure 2b As shown in the embodiment of this application, the first partial image located in the middle of the panoramic image is divided into two sub-images, left and right.
[0065] S134. Perform feature recognition processing on the sub-image to obtain the image features of the sub-image.
[0066] In this embodiment of the application, by feature recognition processing, the scenery, buildings and people in the sub-image can be identified, and by labeling and segmenting the scenery, buildings and people in the sub-image, the background part and the foreground part of the sub-image can be determined, and can be represented in the form of a mask image. Here, the mask image refers to presenting the background and foreground (the foreground may include buildings and people) of the sub-image through images with different gray levels.
[0067] In some embodiments, the method for performing feature recognition processing on sub-images to determine the image features of sub-images includes:
[0068] Determine the sub-images and their shapes;
[0069] Compare the shape of the sub-image with the preset shape:
[0070] When the shape of a sub-image is the same as the preset shape, the sub-image is determined to be the target sub-image;
[0071] When the shape of a sub-image is different from the preset shape, the shape of the sub-image is adjusted to obtain a target sub-image that conforms to the preset shape;
[0072] Feature recognition processing is performed on the target sub-image to obtain its image features;
[0073] Based on the image features of the target sub-image, determine the image features of the sub-image.
[0074] In some embodiments, the method for determining the image features of a sub-image based on a target sub-image includes:
[0075] The target sub-image is encoded and decoded to obtain the mask image of the target sub-image;
[0076] Based on the mask image of the target sub-image, determine the image features of the target sub-image;
[0077] Determine the proportional relationship between the sub-image and the target sub-image;
[0078] The image features of the sub-image are determined based on the proportional relationship between the sub-image and the target sub-image.
[0079] The preset shape refers to the shape that meets the shape requirements of the input image during encoding and decoding processing. For example, in this embodiment, the preset shape is a square. Comparing the shape of the sub-image with the preset shape means determining whether the shape of the sub-image is a square. When the shape of the sub-image is a square, the sub-image is determined to be the target sub-image. When the shape of the sub-image is not a square, the shape of the sub-image is adjusted so that the sub-image conforms to the target sub-image with the preset shape.
[0080] In some embodiments, the method for shape adjustment may include:
[0081] Obtain the size information of the sub-image. The sub-image is typically a rectangular image, and its size information includes its length and width. The size information of the sub-image can be determined by measuring it.
[0082] Based on the size information of the sub-image, the sub-image is resized. This resizing process involves shortening the rectangular image along its length or lengthening it along its width, making the sub-image the same in both length and width. This results in a target sub-image with identical dimensions, which facilitates convolution and pooling processes during feature recognition, leading to better feature identification.
[0083] Encoding and decoding processes refer to feature extraction and upsampling of a target sub-image to obtain a mask image of the target sub-image. Feature extraction can involve convolution and pooling of the target sub-image with the same dimensions. Upsampling can involve deconvolution of the feature-extracted target sub-image to obtain a mask image. The mask image of the target sub-image includes a foreground and a background, where the grayscale values of pixels in the foreground are greater than those in the background, and the foreground represents the image features of the target sub-image.
[0084] The proportional relationship between a sub-image and a target sub-image refers to the transformation relationship between the target sub-image after shape adjustment and the sub-image before shape adjustment. Based on the transformation relationship, the proportional relationship between the sub-image and the target sub-image can be directly determined, and then the image features of the sub-image can be quickly determined based on the image features of the target sub-image.
[0085] S135. Determine the image features of the first local image based on the image features of the sub-image.
[0086] The methods for determining the image features of the first local image include:
[0087] Determine the image features of the sub-image and its position in the first local image;
[0088] The image features of the first local image are determined based on the image features of the sub-image and the position of the sub-image in the first local image.
[0089] Determining the position of a sub-image in the first partial view can refer to determining the coordinate position of the sub-image on the first partial view. For example, in some embodiments, a coordinate system can be established on the first partial view to determine the coordinate position of the sliding window after each sliding step, and then the coordinate position of the sub-image on the first partial view can be determined based on the coordinate position after each sliding step.
[0090] After determining the coordinate position of the sub-image on the first local image, the position of the image features of the sub-image in the first local image is determined by determining the position of the image features of the sub-image in the sub-image, thereby determining the image features of the first local image.
[0091] In some embodiments, when the overlapping region of adjacent sub-images includes overlapping image features, the image features within the overlapping region can be determined by measuring the feature intensity of the overlapping image features within the overlapping region. The feature intensity can be pixel intensity, i.e., the brightness value of a pixel, where the pixel brightness value is between 0 and 255, with values closer to 255 indicating higher brightness and values closer to 0 indicating lower brightness.
[0092] For example, in some embodiments, the sub-image includes a first sub-image and a second sub-image, and the first sub-image and the second sub-image include an overlapping region that overlaps with each other, wherein the image features of the first sub-image in the overlapping region are first image features, and the image features of the second sub-image in the overlapping region are second image features;
[0093] The method for obtaining image features of the first local image by performing feature recognition processing on the first local image also includes:
[0094] When the first image feature overlaps with the second image feature, calculate the first feature intensity of the first image feature and the second feature intensity of the second image feature;
[0095] Compare the intensity of the first feature and the intensity of the second feature:
[0096] When the intensity of the first feature is greater than the intensity of the second feature, the first image feature is selected as the image feature of the first local image;
[0097] When the intensity of the first feature is less than the intensity of the second feature, the second image feature is selected as the image feature of the first local image.
[0098] When the intensity of the first feature is equal to the intensity of the second feature, the first image feature or the second image feature is selected as the image feature of the first local image.
[0099] S140. Determine the target feature element based on multiple feature elements, and determine the position of the target feature element in the first local graph.
[0100] The location of the target feature element in the first local image can be determined by extracting the feature elements. This extraction process involves identifying the most significant feature elements based on their pixel values. A significant feature element is defined as a target feature element whose pixel value is greater than a preset pixel threshold. The pixel value can be the pixel intensity, ranging from 0 to 255. The pixel threshold can be manually set; for example, it can be preset to 150 or 200 as needed. The target feature element is thus defined as a feature element with a pixel value greater than 150 or 200.
[0101] like Figure 4 As shown, in some embodiments, the method for determining a target feature element based on multiple feature elements and determining the position of the target feature element in the first local image includes:
[0102] S141. Determine the pixel values of multiple feature elements in the first local image.
[0103] Feature elements are the pixels that make up the image features. The pixel value of each pixel can be determined, and the pixel value can be the grayscale value of the pixel. Methods for determining the pixel values of feature elements in the first local image can include obtaining the grayscale value of the pixel using the impixel function.
[0104] S142. Compare the pixel values of multiple feature elements in the first local image with a preset pixel threshold, and determine the first feature element whose pixel value is greater than the pixel threshold and the second feature element whose pixel value is less than the pixel threshold, wherein the first feature element is the target feature element.
[0105] S143. Binarize the pixel values of the first feature element and the second feature element to obtain the first local image after binarization.
[0106] Binarization refers to increasing the pixel value of the first feature element and decreasing the pixel value of the second feature element based on the intensity comparison result, making the difference between the first and second feature elements more obvious. For example, in some embodiments, the grayscale value of a pixel is between 0 and 255. Increasing the pixel value of the first feature element can mean adjusting the pixel value of the first feature element to 255, and decreasing the pixel value of the second feature element can mean adjusting the pixel value of the second feature element to 0.
[0107] S144. Based on the first local image after binarization, determine the position of the target feature element in the first local image.
[0108] The position of the target feature element in the first local graph can be determined by the positional relationship between the target feature element and the first local graph.
[0109] S150. Based on the preset segmentation model, determine the positional relationship between the first local image and the panoramic image, as well as the position of the target feature element in the first local image, and determine the position of the target feature element in the panoramic image.
[0110] Determining the positional relationship between the first local image and the panoramic image refers to determining the position of the first local image in the panoramic image based on the size and position of the panoramic image when segmenting it according to the preset segmentation model. Since the preset segmentation model is pre-set, it can be directly and quickly obtained based on the corresponding positional change relationship when determining the positional relationship.
[0111] Determining the position of the target feature element in the panoramic image means that since the positional relationship between the first local image and the panoramic image is determined, the position of the target feature element in the first local image is also determined. Therefore, the position of the target feature element in the panoramic image can be quickly and accurately determined by coordinate conversion.
[0112] S160. Based on the position of the target feature elements in the panoramic image, perform distortion correction processing on the panoramic image to obtain the corrected panoramic image.
[0113] Distortion correction processing refers to adjusting the distance between the lens and the imaging plane during image acquisition to achieve a sharper image of the subject. Essentially, it uses the location of the target feature element in the panoramic image as the correction point, ensuring the image at that location is sharpest and has minimal distortion. Correction methods can include changing the focus point to align with the target feature element. Focusing refers to adjusting the distance between the lens and the imaging plane during image acquisition to achieve a sharper image of the subject. The focus point is the point on the subject represented by the desired feature element.
[0114] like Figure 5 As shown, after obtaining the corrected panoramic image, the image processing method of this application further includes:
[0115] S170. Obtain the position of the target feature element in the modified panoramic image;
[0116] S180. Based on the position of the target feature element in the corrected panoramic image, obtain the next frame panoramic image, and perform distortion correction processing on the next frame panoramic image to obtain the corrected next frame panoramic image.
[0117] The next frame panoramic image can be the next frame of a panoramic video or the next frame of a continuously captured panoramic image. The method for obtaining the next frame panoramic image can be through capturing the image itself or by retrieving a pre-stored panoramic image. Distortion correction processing of the obtained next frame panoramic image involves capturing the next frame panoramic image based on the location of the target feature elements in the modified currently captured panoramic image, and then using the target feature elements of the modified panoramic image as the clearest and least distorted location in the next frame panoramic image for distortion correction processing, resulting in the corrected next frame panoramic image.
[0118] The image processing method in this embodiment of the invention will be described below with reference to a specific application scenario.
[0119] Please see Figure 6 This is a schematic diagram illustrating an example of the image processing method applied in an experimental scenario according to an embodiment of the present invention. The image processing method is applied to a server and includes:
[0120] S210, Obtain panoramic image.
[0121] The horizontal field of view of the panoramic image is 0–360°, and the vertical field of view is 0–180°.
[0122] S220. The panoramic image is segmented to obtain a first local image and a second local image. The distortion degree of the first local image is lower than that of the second local image.
[0123] The panoramic image is divided into upper and lower sections, each with a 10% area. This 10% area is used to create a second partial image, while the remaining 80% is used to create a first partial image. For example, if the panoramic image is 1000 pixels long and 500 pixels wide, the first partial image after cropping will have a length of 1000 pixels and a width of 400 pixels.
[0124] S230. Extract and process the first local image to obtain a sub-image.
[0125] The first partial image is extracted using a sliding window to obtain two sub-images, with an overlapping area between them. For example, the first sub-image has a length of 600 and a width of 400, and the second sub-image has a length of 600 and a width of 400. The overlapping area between the first and second sub-images has a length of 200 and a width of 400.
[0126] S240. Perform feature recognition processing on the sub-image to obtain the image features of the sub-image.
[0127] U 2 The net performs feature recognition processing on the sub-image. Specifically, the sub-image is input into the encoder-decoder network structure model to obtain 6 mask images with the same size as the input sub-image. The intensity of the 6 mask images is averaged and output to obtain the sub-image mask.
[0128] S250. Determine the position of the sub-image in the first local image, and determine the image features of the first local image based on the position of the sub-image in the first local image. The image features of the first local image are composed of multiple feature elements.
[0129] A coordinate system is established on the first local image, where the coordinates of the four corners of the first local image are (0, 0), (1000, 0), (1000, 400), and (0, 400). Therefore, the coordinates of the four corners of the first sub-image are (0, 0), (600, 0), (600, 400), and (0, 400). The coordinates of the four corners of the second sub-image are (400, 0), (1000, 0), (1000, 400), and (400, 400). Based on the coordinate relationship between the sub-image and the first local image, the positional relationship between the sub-image and the first local image can be determined. Furthermore, based on the positional relationship between the mask image of the sub-image and the first local image, a mask image used to characterize the image features of the first local image is determined.
[0130] S260. Extract the feature elements of the image features of the first local image, determine the target feature element among the feature elements, and determine the position of the target feature element in the first local image.
[0131] The pixel value of each pixel in the mask image used to represent the image features of the first local image is determined. The mask image is then binarized, adjusting the pixel values of all pixels with intensity values greater than a preset threshold to 255, and adjusting the pixel values of all pixels with intensity values less than the preset threshold to 0. This yields a binary image of the first local image, facilitating the identification and selection of target feature elements and determining their positions within the first local image.
[0132] S270. Determine the position of the target feature element in the panoramic image based on the positional relationship between the first local image and the panoramic image.
[0133] The positional relationship between the first local image and the panoramic image is determined based on the segmentation ratio. Since the segmentation ratio between the first local image and the panoramic image is determined during segmentation processing, the position of the target feature element in the panoramic image can be quickly determined.
[0134] S280. Based on the position of the target feature elements in the panoramic image, perform distortion correction processing on the panoramic image to obtain the corrected panoramic image.
[0135] In this embodiment, by segmenting the panoramic image, a first local image with lower distortion is quickly identified. Feature recognition and extraction are then performed on the first local image to determine its target feature elements. Based on the segmentation ratio, the position of these target feature elements within the panoramic image is quickly determined. Finally, distortion correction is applied to the panoramic image based on these positions, resulting in a corrected panoramic image. This improves the processing speed of panoramic images and overcomes the problem of cumbersome and slow processing in existing methods.
[0136] To better implement the above methods, this application also provides an image processing apparatus, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.
[0137] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the image processing device specifically integrated into the server as an example.
[0138] For example, such as Figure 7 As shown, the image processing apparatus may include:
[0139] Acquisition unit 301 is used to acquire panoramic images;
[0140] The segmentation processing unit 302 is used to segment the panoramic image using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image.
[0141] The feature recognition processing unit 303 is used to determine a target feature element based on a plurality of feature elements, and to determine the position of the target feature element in the first local image;
[0142] The first determining unit 304 is used to determine the positional relationship between the first local image and the panoramic image, as well as the position of the target feature element in the first local image, and the position of the target feature element in the panoramic image, according to the preset segmentation model.
[0143] The second determining unit 305 is used to determine the positional relationship between the first local image and the panoramic image according to a preset segmentation model; and to determine the position of the target feature element in the panoramic image according to the positional relationship between the first local image and the panoramic image, and the position of the target feature element in the first local image.
[0144] The correction unit 306 is used to perform distortion correction processing on the panoramic image based on the position of the target feature elements in the panoramic image to obtain the corrected panoramic image.
[0145] In some embodiments of this application, the feature recognition processing unit 303 is further specifically used for:
[0146] Determine the dimensions of the first partial view;
[0147] Based on the dimensional information in the first partial view, determine the dimensions, sliding step size, and sliding direction of the sliding window;
[0148] Based on the size, sliding step, and sliding direction of the sliding window, the first local image is segmented to obtain multiple sub-images;
[0149] The sub-images are processed by feature recognition to obtain their image features;
[0150] Based on the image features of the sub-images, determine the image features of the first local image.
[0151] In some embodiments of this application, the feature recognition processing unit 303 is further specifically used for:
[0152] Determine the image features of the sub-image and its position in the first local image;
[0153] The image features of the first local image are determined based on the image features of the sub-image and the position of the sub-image in the first local image.
[0154] In some embodiments of this application, the feature recognition processing unit 303 is further specifically used for:
[0155] Determine the sub-images and their shapes;
[0156] Compare the shape of the sub-image with the preset shape:
[0157] When the shape of a sub-image is the same as the preset shape, the sub-image is determined as the target sub-image; when the shape of a sub-image is different from the preset shape, the shape of the sub-image is adjusted to obtain the target sub-image that conforms to the preset shape.
[0158] Feature recognition processing is performed on the target sub-image to obtain its image features;
[0159] Based on the image features of the target sub-image, determine the image features of the sub-image.
[0160] In some embodiments of this application, the feature recognition processing unit 303 is further specifically used for:
[0161] The target sub-image is encoded and decoded to obtain the mask image of the target sub-image;
[0162] Based on the mask image of the target sub-image, determine the image features of the target sub-image;
[0163] Determine the proportional relationship between the sub-image and the target sub-image;
[0164] Based on the proportional relationship between the sub-image and the target sub-image, the image features of the sub-image are determined.
[0165] In some embodiments of this application, the feature recognition processing unit 303 is further specifically used for:
[0166] The sub-image includes a first sub-image and a second sub-image, and there is an overlapping region between the first sub-image and the second sub-image. The image features of the first sub-image in the overlapping region are the first image features, and the image features of the second sub-image in the overlapping region are the second image features.
[0167] The method for obtaining image features of the first local image by performing feature recognition processing on the first local image also includes:
[0168] When the first image feature overlaps with the second image feature, calculate the first feature intensity of the first image feature and the second feature intensity of the second image feature;
[0169] Compare the intensity of the first feature and the intensity of the second feature:
[0170] When the intensity of the first feature is greater than the intensity of the second feature, the first image feature is selected as the image feature of the first local image;
[0171] When the intensity of the first feature is less than the intensity of the second feature, the second image feature is selected as the image feature of the first local image.
[0172] When the intensity of the first feature is equal to the intensity of the second feature, the first image feature or the second image feature is selected as the image feature of the first local image.
[0173] In some embodiments of this application, the first determining unit 304 is further specifically used for:
[0174] Determine the pixel values of multiple feature elements in the first local image;
[0175] The pixel values of multiple feature elements in the first local image are compared with a preset pixel threshold, and a first feature element whose pixel value is greater than the pixel threshold and a second feature element whose pixel value is less than the pixel threshold are determined respectively, wherein the first feature element is the target feature element.
[0176] The pixel values of the first feature element and the second feature element are binarized to obtain the first local image after binarization.
[0177] Based on the first local image after binarization, determine the position of the target feature element in the first local image.
[0178] In some embodiments of this application, the correction unit 306 is further specifically used for:
[0179] Obtain the position of the target feature elements in the corrected panoramic image;
[0180] Based on the position of the target feature element in the corrected panoramic image, the next frame of the panoramic image is obtained, and distortion correction processing is performed on the next frame of the panoramic image to obtain the corrected next frame of the panoramic image.
[0181] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0182] As can be seen from the above, the image processing device of this embodiment comprises: an acquisition unit 301 for acquiring a panoramic image; a segmentation processing unit 302 for segmenting the panoramic image using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image; a feature recognition processing unit 303 for performing feature recognition processing on the first local image to obtain image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements; a first determination unit 304 for determining a target feature element based on the multiple feature elements and determining the position of the target feature element in the first local image; a second determination unit 305 for determining the positional relationship between the first local image and the panoramic image, and the position of the target feature element in the first local image, and determining the position of the target feature element in the panoramic image, based on the preset segmentation model; and a correction unit 306 for performing distortion correction processing on the panoramic image based on the position of the target feature element in the panoramic image to obtain a corrected panoramic image. Therefore, this embodiment can improve the speed of processing panoramic images and address the problem of the cumbersome and slow processing speed of existing recognition processes.
[0183] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0184] In some embodiments, the image processing apparatus may also be integrated into multiple electronic devices, such as multiple servers, with the image processing method of this application being implemented by multiple servers.
[0185] In this embodiment, the electronic device will be described in detail as an image processing device, for example, such as... Figure 8 As shown, it illustrates a structural schematic diagram of the image processing apparatus involved in the embodiments of this application. Specifically:
[0186] The image processing device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will understand that... Figure 8 The image processing apparatus structure shown does not constitute a limitation on the image processing apparatus. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0187] The processor 401 is the control center of the image processing device, connecting various parts of the device via various interfaces and lines. It executes various functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0188] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the image processing device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0189] The image processing apparatus also includes a power supply 403 that supplies power to the various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0190] The image processing device may also include an input module 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0191] The image processing device may also include a communication module 405. In some embodiments, the communication module 405 may include a wireless module, through which the image processing device can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media.
[0192] Although not shown, the image processing apparatus may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the image processing apparatus loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions.
[0193] In some embodiments, a computer program product is also provided, comprising a computer program or instructions that, when executed by a processor, implement the steps in any of the above-described image processing methods.
[0194] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0195] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0196] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the image processing methods provided in embodiments of this application.
[0197] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0198] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the image processing aspects described in the above embodiments.
[0199] Since the instructions stored in the storage medium can execute the steps of any of the image processing methods provided in the embodiments of this application, the beneficial effects that any of the image processing methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0200] The foregoing has provided a detailed description of an image processing method, apparatus, terminal, storage medium, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An image processing method, characterized by, include: Acquire panoramic images; The panoramic image is segmented using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image. The first local image is subjected to feature recognition processing to obtain the image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements; A target feature element is determined based on multiple feature elements, and the position of the target feature element in the first local image is determined. Based on the preset segmentation model, the positional relationship between the first local image and the panoramic image is determined, as well as the position of the target feature element in the first local image, and the position of the target feature element in the panoramic image is determined. The position of the target feature element in the panoramic image is used as the position for distortion correction processing to correct the panoramic image, resulting in a corrected panoramic image.
2. The image processing method of claim 1, wherein, After obtaining the corrected panoramic image, the method further includes: Obtain the position of the target feature element in the corrected panoramic image; Based on the position of the target feature element in the corrected panoramic image, the next frame of the panoramic image is obtained, and distortion correction processing is performed on the next frame of the panoramic image to obtain the corrected next frame of the panoramic image.
3. The image processing method according to claim 1, characterized in that, The step of performing feature recognition processing on the first local image to obtain the image features of the first local image includes: Determine the size information of the first partial image; Based on the size information of the first partial image, determine the size, sliding step size, and sliding direction of the sliding window; Based on the size, sliding step size, and sliding direction of the sliding window, the first local image is segmented to obtain multiple sub-images; The sub-image is subjected to feature recognition processing to obtain the image features of the sub-image; Based on the image features of the sub-image, the image features of the first local image are determined.
4. The image processing method according to claim 3, characterized in that, Determining the image features of the first local image based on the image features of the sub-image includes: Determine the image features of the sub-image and the position of the sub-image in the first local image; The image features of the first local image are determined based on the image features of the sub-image and the position of the sub-image in the first local image.
5. The image processing method according to claim 3, characterized in that, The step of performing feature recognition processing on the sub-image to determine the image features of the sub-image includes: Determine the sub-image and the shape of the sub-image; The shape of the sub-image is compared with a preset shape: When the shape of the sub-image is the same as the preset shape, the sub-image is determined to be the target sub-image; When the shape of the sub-image is different from the preset shape, the shape of the sub-image is adjusted to obtain a target sub-image that conforms to the preset shape; The target sub-image is subjected to feature recognition processing to obtain the image features of the target sub-image; The image features of the sub-image are determined based on the image features of the target sub-image.
6. The image processing method according to claim 5, characterized in that, The method for determining the image features of the sub-image includes: The target sub-image is encoded and decoded to obtain a mask image of the target sub-image; Based on the mask image of the target sub-image, determine the image features of the target sub-image; Determine the proportional relationship between the sub-image and the target sub-image; The image features of the sub-image are determined based on the proportional relationship between the sub-image and the target sub-image.
7. The image processing method according to claim 3, characterized in that, The sub-image includes a first sub-image and a second sub-image, and the first sub-image and the second sub-image include an overlapping region that overlaps with each other, wherein the image features of the first sub-image in the overlapping region are first image features, and the image features of the second sub-image in the overlapping region are second image features; The method for performing feature recognition processing on the first local image to obtain the image features of the first local image further includes: When the first image feature overlaps with the second image feature, calculate the first feature intensity of the first image feature and the second feature intensity of the second image feature; Compare the first feature intensity and the second feature intensity: When the intensity of the first feature is greater than the intensity of the second feature, the first image feature is selected as the image feature of the first local image; When the intensity of the first feature is less than the intensity of the second feature, the second image feature is selected as the image feature of the first local image; When the first feature intensity is equal to the second feature intensity, the first image feature or the second image feature is selected as the image feature of the first local image.
8. The image processing method according to claim 1, characterized in that, The step of determining a target feature element based on a plurality of feature elements and determining the position of the target feature element in the first local image includes: Determine the pixel values of multiple feature elements in the first local image; The pixel values of multiple feature elements in the first partial image are compared with a preset pixel threshold, and a first feature element whose pixel value is greater than the pixel threshold and a second feature element whose pixel value is less than the pixel threshold are determined respectively, wherein the first feature element is a target feature element. The pixel values of the first feature element and the pixel values of the second feature element are binarized to obtain the first local image after binarization. The position of the target feature element in the first local image is determined based on the first local image after binarization.
9. An image processing apparatus, characterized in that, include: The acquisition unit is used to acquire panoramic images; The segmentation processing unit is used to segment the panoramic image using a preset segmentation model to obtain a first local image and a second local image, wherein the distortion degree of the first local image is lower than that of the second local image. The feature recognition processing unit is used to perform feature recognition processing on the first local image to obtain the image features of the first local image, wherein the image features of the first local image are composed of multiple feature elements; The first determining unit is configured to determine a target feature element based on the plurality of feature elements, and to determine the position of the target feature element in the first local image; The second determining unit is used to determine the positional relationship between the first local image and the panoramic image, as well as the position of the target feature element in the first local image, and the position of the target feature element in the panoramic image, based on the preset segmentation model. The correction unit is used to use the position of the target feature element in the panoramic image as the position for distortion correction processing to correct the panoramic image, thereby obtaining a corrected panoramic image.
10. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the image processing method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the image processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Self-adaptive segmentation and distortion-free expansion system and method for panoramic annular image
CN110458753A
Image processing method and device, electronic equipment and storage medium
CN111091507A