Image processing method, device and equipment and computer readable storage medium

The prediction model preprocesses and feature extraction of fisheye images, and directly realizes the identification and positioning of target objects, solving the problem of slow computing speed after distortion correction of fisheye images in the prior art, and improving efficiency.

CN120220176APending Publication Date: 2025-06-27GUANGDONG JUHUA RES INST OF ADVANCED DISPLAY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311800457.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

After the fish eye image is distorted and corrected, the prior art uses large amount of calculation and slow calculation speed to reduce efficiency.

Method used

The fisheye image is preprocessed through the prediction model, the target feature information is extracted, and the target position information is obtained through the second processing step, which directly realizes the recognition and positioning of the target object, avoids the distortion correction step.

Benefits of technology

The efficiency of target object recognition and positioning based on fisheye images is improved, and the calculation amount and calculation time are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220176A_ABST
    Figure CN120220176A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method, apparatus and device, and a computer readable storage medium. The method comprises the steps of obtaining to-be-processed image information; performing first processing on the to-be-processed image information through a prediction model to obtain target feature information; and performing second processing on the target feature information to obtain target position information. The to-be-processed image information is the fisheye image, the fisheye image is predicted through the pre-trained prediction model, the target position information of the target object in the fisheye image is obtained, and compared with an existing method that distortion correction needs to be carried out on the fisheye image, and the corrected image is obtained, the image processing efficiency is improved. And then the target object is identified and positioned based on the corrected image, so that the efficiency of identifying and positioning the target object based on the fisheye image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to an image processing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of AI technology, the concept of the metaverse has gradually matured, and the AR / VR field has gradually come into the public eye. At the same time, the fisheye camera has a large viewing angle range and can present more information. For the convenience of human-computer interaction, AR / VR device manufacturers often use fisheye cameras to obtain fisheye images of target objects, and then realize the recognition and positioning of target objects.

[0003] Existing methods need to correct the distortion of the fisheye image to obtain the corrected image, and then perform the recognition and positioning of the target object based on the corrected image. However, the distortion correction of the fisheye image requires a large amount of image matrix calculations, resulting in a large amount of calculation, slow operation speed, and reduced efficiency of recognizing and positioning the target object based on the fisheye image. Therefore, how to improve the efficiency of recognizing and positioning the target object based on the fisheye image is an urgent problem to be solved. Summary of the Invention

[0004] Embodiments of this application provide an image processing method, apparatus, device, and computer-readable storage medium, which can improve the efficiency of recognizing and positioning a target object based on a fisheye image.

[0005] In a first aspect, embodiments of this application provide an image processing method, and the method includes:

[0006] Obtain information of an image to be processed;

[0007] Perform a first processing on the information of the image to be processed through a prediction model to obtain target feature information;

[0008] Perform a second processing on the target feature information to obtain target position information.

[0009] In a second aspect, embodiments of this application provide an image processing apparatus, including:

[0010] An obtaining unit, configured to obtain information of an image to be processed;

[0011] A first processing unit, configured to perform a first processing on the information of the image to be processed through a prediction model to obtain target feature information;

[0012] A second processing unit, configured to perform a second processing on the target feature information to obtain target position information.

[0013] In a third aspect, an embodiment of the present application further provides an image processing device, including a memory storing multiple computer programs; a processor loads the computer programs from the memory to execute the steps of any one of the image processing methods provided by the embodiments of the present application.

[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing multiple computer programs, and the computer programs are suitable for being loaded by a processor to execute the steps of any one of the image processing methods provided by the embodiments of the present application.

[0015] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps in any one of the image processing methods provided by the embodiments of the present application are implemented.

[0016] Adopting the solution of the embodiment of the application, obtain the image information to be processed; perform a first process on the image information to be processed through a prediction model to obtain target feature information; perform a second process on the target feature information to obtain target position information. The image information to be processed is a fish-eye image. By predicting the fish-eye image through a pre-trained prediction model, the target position information of the target object in the fish-eye image is obtained. Compared with the existing method that requires distortion correction of the fish-eye image, obtaining the corrected image, and then performing recognition and positioning of the target object based on the corrected image, the efficiency of recognizing and positioning the target object based on the fish-eye image is improved. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the first embodiment of the image processing method provided by the embodiment of the present application;

[0019] Figure 2 It is a schematic flowchart of the second embodiment of the image processing method provided by the embodiment of the present application;

[0020] Figure 3 It is a schematic structural diagram of the image processing device provided by the embodiment of the present application;

[0021] Figure 4 It is a schematic structural diagram of the image processing device provided by the embodiment of the present application. Detailed Embodiments

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0023] The embodiments of the present application provide an image processing method, apparatus, device, and computer-readable storage medium.

[0024] Specifically, the embodiments of the present application will be described from the perspective of an image processing apparatus, which can be specifically integrated in an image processing device, that is, the image processing method in the embodiments of the present application can be executed by the image processing device.

[0025] The image processing method provided by the embodiments of the present application can be applied to an image processing device, and the image processing device can include devices such as smart terminals, PC terminals, AR / VR devices, etc.

[0026] The following will be described in detail with reference to the accompanying drawings. In the embodiments of the present application, the execution entity is an image processing device as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown in the drawings.

[0027] Please refer to Figure 1 , the specific process of the first embodiment of the image processing method includes the following steps:

[0028] Step 101, obtain the image information to be processed;

[0029] Step 102, perform a first process on the image information to be processed through a prediction model to obtain target feature information;

[0030] Step 103, perform a second process on the target feature information to obtain target position information.

[0031] In this embodiment, during the process of a user using an image processing device, the image processing device acquires a fisheye image of a target object and historical position information of the target object through a fisheye camera, performs a first processing on the fisheye image through a prediction model to obtain target feature information corresponding to the target object in the fisheye image, and performs a second processing on the target feature information and the historical position information through the prediction model to obtain target position information of the target object, thereby realizing the recognition and positioning of the target object in the fisheye image. It should be noted that the image processing device may be an AR / VR device worn in front of the user's eyes, and the user controls the image processing device to execute corresponding functions through hand movements. Therefore, the target object is usually the user's hand, and the image processing device needs to recognize and position the user's hand to facilitate subsequent recognition of the user's hand movements. Among them, the historical position information of the target object is the position information corresponding to the target object in the previous fisheye image of the current frame fisheye image acquired by the image processing device; the prediction model is pre-trained to be able to recognize and position the target object in the fisheye image.

[0032] The image processing device in this embodiment acquires a fisheye image and historical position information of a target object in the fisheye image; performs a first processing on the fisheye image through a prediction model to obtain target feature information corresponding to the fisheye image; and performs a second processing on the target feature information and the historical position information through the prediction model to obtain target position information of the target object. By predicting the fisheye image through a pre-trained prediction model to obtain the target position information of the target object in the fisheye image, compared with the existing method that requires distortion correction of the fisheye image, obtaining the corrected image, and then performing recognition and positioning of the target object based on the corrected image, the efficiency of recognizing and positioning the target object based on the fisheye image is improved.

[0033] Specifically, each step will be described in detail as follows:

[0034] Step 101, acquire image information to be processed;

[0035] In this step, during the process of a user using an image processing device, the image processing device acquires image information to be processed through a fisheye camera, and the image information to be processed includes a fisheye image of a target object. Exemplarily, the image processing device needs to recognize and position the user's hand to facilitate subsequent recognition of the user's hand movements. Therefore, the target object is usually the user's hand, and the image processing device acquires a fisheye image containing the user's hand through the fisheye camera.

[0036] Step 102, perform a first processing on the image information to be processed through a prediction model to obtain target feature information;

[0037] In this step, after obtaining the image information to be processed, the image processing device inputs the image information to be processed into a pre-trained prediction model. The prediction model performs a first processing on the fish-eye image to obtain target feature information corresponding to the image information to be processed. The image information to be processed includes the fish-eye image of the target object. The first processing includes convolution processing, pooling processing, and normalization processing. The prediction model sequentially performs convolution processing and pooling processing on the fish-eye image, extracts shape features, color features, texture features, and background features in the fish-eye image, and then performs normalization processing on the extracted features to obtain the target feature information corresponding to the fish-eye image. It can be understood that the shape features, color features, and texture features all belong to the features of the target object in the fish-eye image, and the background features belong to the features of the non-target object area in the fish-eye image. Exemplarily, the image processing device needs to identify and locate the user's hand to facilitate subsequent recognition of the user's hand movements. Therefore, the target object is usually the user's hand. The shape features, color features, and texture features in the obtained target feature information all belong to the features of the user's hand in the fish-eye image, and the background features belong to the features of the non-user's hand area in the fish-eye image.

[0038] Specifically, step 102 includes:

[0039] Step 1021, performing convolution processing and pooling processing on the image information to be processed through the prediction model to obtain first feature information;

[0040] In this step, the image processing device performs convolution processing and pooling processing on the image information to be processed through the prediction model to obtain first feature information. Specifically, the image information to be processed includes the fish-eye image of the target object. The first feature information is downsampled feature information. The image processing device first performs convolution processing on the fish-eye image through the prediction model to obtain shape features, color features, texture features, and background features in the fish-eye image, and then performs pooling processing on the shape features, color features, texture features, and background features in the fish-eye image to obtain downsampled feature information corresponding to the shape features, color features, texture features, and background features in the fish-eye image. It can be understood that the prediction model may include one or more convolutional layers. Each convolutional layer includes multiple convolutional kernels. Each convolutional kernel is responsible for extracting a certain feature in the fish-eye image. For example, if it is necessary to extract shape features, color features, texture features, and background features in the fish-eye image, each convolutional layer includes four convolutional kernels, and each convolutional kernel corresponds to extracting a certain feature. The prediction model may include one or more pooling layers. A pooling layer can be set separately after each convolutional layer, or a pooling layer can be set for multiple convolutional layers. Preferably, a pooling layer is set for multiple convolutional layers, which reduces the dimension of the features obtained by the convolutional layer, speeds up convergence, is beneficial to selecting more stable and better-invariant features, and improves the generalization ability at the same time.

[0041] Further, step 1021 includes:

[0042] Step 10211, performing convolution processing on the to-be-processed image information through a prediction model to obtain semantic feature information;

[0043] In this step, the image processing device performs convolution processing on the to-be-processed image information through a prediction model to obtain semantic feature information. Specifically, the to-be-processed image information includes a fisheye image of a target object. The processing device performs convolution processing on the fisheye image through the prediction model based on the number of convolution kernels and the target step size, and obtains the semantic feature information corresponding to the fisheye image. Among them, the semantic feature information includes a set of feature maps corresponding to shape features, color features, texture features, and background features. Specifically, the prediction model includes multiple convolutional layers, and each convolutional layer includes multiple convolution kernels. The convolution kernel is a three-dimensional convolution kernel. In each convolutional layer, the prediction model sequentially performs convolution processing on the input data of each convolutional layer through each convolution kernel according to the target step size, and then uses the output data of each convolutional layer as the input data of the next convolutional layer. The prediction model performs layer-by-layer feature extraction through multiple convolutional layers until the last convolutional layer, and respectively obtains a set of feature maps corresponding to shape features, color features, texture features, and background features.

[0044] Exemplarily, the fisheye image is a fisheye image including a user's hand, and the target object is the user's hand. Assume that the size of the convolution kernel is (K x ,K y ,K z ), there are four convolution kernels in the convolutional layer, and the target step sizes are set as S x and S y ; in the process of performing convolution in the first convolutional layer, the input is a fisheye image including a user's hand. The prediction model respectively makes the four convolution kernels perform convolution in the x and y directions of the fisheye image including the user's hand according to the target step sizes S x and S y , and respectively obtains a set of feature maps corresponding to shape features, color features, texture features, and background features; in the process of performing convolution in the second convolutional layer, the input is a set of feature maps corresponding to shape features, color features, texture features, and background features. The prediction model respectively makes the four convolution kernels perform convolution in the x and y directions of each feature map according to the target step sizes S x and S y , and further obtains a set of feature maps corresponding to deeper shape features, color features, texture features, and background features; the convolution process of subsequent convolutional layers refers to the convolution process of the second convolutional layer, and finally obtains a set of feature maps corresponding to shape features, color features, texture features, and background features with the target depth.

[0045] It should be noted that the output of each convolutional layer includes M n feature maps of size , where M n is determined by the number of convolutional kernels in the current layer. M refers to the number of required convolutional kernels, n refers to the layer number, and are as follows:

[0046]

[0047] where n is the label of the layer, is the size of the output feature map.

[0048] Among the M n feature maps output by the convolutional layer, the data vector at the position (i, j) of each feature map is:

[0049] where x ik is the data vector at the position (i, j) in the feature map output by the previous layer, K = {(K x , K y , K z ), S = {(S x , S y ), f kS is the convolution operation, and δi and δj are target parameters determined during the training of the prediction model.

[0050] Step 10212: Perform pooling processing on each semantic feature in the semantic feature information through the prediction model to obtain first feature information.

[0051] In this step, the image processing device performs pooling processing on each semantic feature in the semantic feature information through the prediction model to obtain first feature information. Specifically, the image information to be processed includes the fisheye image of the target object. The first feature information is downsampled feature information. The image processing device performs convolution processing on the fisheye image through the prediction model to obtain semantic feature information including a set of feature maps corresponding to shape features, color features, texture features, and background features. Then, for each feature map in the set of feature maps in the semantic feature information, the image processing device performs pooling processing on each feature map through the prediction model based on the target pooling size to obtain the downsampled feature information corresponding to the fisheye image. It can be understood that the image processing device performs convolution processing on the fisheye image through the prediction model to obtain multiple feature maps, and each feature map needs to be input to the pooling layer for pooling processing; exemplarily, for the feature map input to the pooling layer by the prediction model, the target pooling size is (P x , P y ), and the pooling layer takes non-overlapping (P x , P y) a square area of size, using the maximum activation function and the target pooling size of (P x , P y ) to perform pooling on the feature map. The pooling layer usually performs a sliding window operation on the feature map output by the convolutional layer, and processes the pixel values within each window as a whole. Usually, the pooling layer will select a window of a fixed size (P x , P y ), (P x , P y ) can be, for example, 2x2 or 3x3, and then perform some simple operations on the pixel values within each window, such as taking the average or the maximum value, to obtain a new pixel value to represent the feature information within the window; the pooling layer can obtain position invariance within a larger local area and achieve downsampling in the x and y directions based on factors P x and P y , and finally obtain the downsampled feature map corresponding to each feature map, and obtain the downsampled feature information corresponding to the fish-eye image.

[0052] Step 1022, normalize the first feature information through the prediction model to obtain the target feature information.

[0053] In this step, the image processing device normalizes the first feature information through the prediction model to obtain the target feature information. Specifically, the first feature information is the downsampled feature information. After the image processing device obtains the downsampled feature information corresponding to the fish-eye image through the prediction model, in order to solve the problems of gradient disappearance and gradient explosion of the data, the prediction model normalizes the downsampled feature information to obtain the target feature information corresponding to the fish-eye image, which can improve the operation speed of the model and reduce the operation cost. Among them, the target feature information is a set of downsampled feature maps corresponding to the shape feature, color feature, texture feature, and background feature after normalization processing.

[0054] Further, step 1022 includes:

[0055] Step 10221, calculate the mean and variance corresponding to the first feature information through the prediction model;

[0056] In this step, the image processing device calculates the mean and variance corresponding to the first feature information through the prediction model; specifically, the first feature information is the downsampled feature information, and each downsampled feature map in the downsampled feature information is represented in the form of a data vector. The prediction model can calculate the mean and variance of the set of downsampled feature maps corresponding to the downsampled feature information according to the number of downsampled feature maps and the data vector in the downsampled feature information.

[0057] Step 10222: Based on the mean value, the variance, and the target weights, perform normalization processing on each first feature in the first feature information through a prediction model to obtain target feature information.

[0058] In this step, the image processing device performs normalization processing on each first feature in the first feature information through a prediction model based on the mean value, the variance, and the target weights to obtain target feature information. Specifically, for each downsampled feature map in the downsampled feature information, the image processing device performs normalization processing on each downsampled feature map through a prediction model based on the mean value, the variance, and the target weights to obtain the target feature information corresponding to the fisheye image. Among them, the target weights are determined during the training process of the prediction model. Specifically, first, the normalization layer of the prediction model calculates the mean value μ and the variance σ of the input set of downsampled feature maps. 2 , for each downsampled feature map in the set of downsampled feature maps, the normalization layer of the prediction model uses μ and σ 2 to perform a normalization operation on the downsampled feature map:

[0059]

[0060] where x i is the data vector of the downsampled feature map, is the intermediate normalized data vector of the data vector of the downsampled feature map, and ε is a very small constant.

[0061]

[0062] where y i is the target normalized data vector of the data vector of the downsampled feature map, and γ and β are the target weights.

[0063] Step 103: Perform a second process on the target feature information to obtain target position information.

[0064] In this step, after obtaining the target feature information, the image processing device performs a second processing on the target feature information to obtain the target position information. Specifically, the image information to be processed includes a fisheye image of the target object. The image processing device first obtains the historical position information of the target object in the fisheye image in the previous frame of the fisheye image from a pre-created historical database, and then performs a second processing on the target feature information and the historical position information through a prediction model to obtain the target position information of the target object. Specifically, the prediction model includes a conversion layer, a fully connected layer, and a multi-layer perceptron layer. The target feature information and the historical position information are input into the conversion layer, and the feature encoding of the target feature information is enhanced by using the attention mechanism in the conversion layer. The deep and shallow target feature information are fused to obtain a Gaussian heat map set, and then the classification head and regression head constructed by the fully connected layer and the multi-layer perceptron layer respectively predict the confidence and position information of the target object corresponding to each Gaussian heat map in the output Gaussian heat map set, and then determine the target confidence in the confidence of the target object corresponding to all Gaussian heat maps, and then use the position information corresponding to the Gaussian heat map corresponding to the target confidence as the target position information of the target object in the fisheye image.

[0065] Among them, the conversion layer includes the Transformer model, which includes N encoders and N decoders. Each encoder layer consists of 1 self-attention layer, 2 normalization layers and 1 feedforward neural network layer; each decoder layer consists of 2 self-attention layers, 3 normalization layers and 1 feedforward neural network layer. Transformer is a landmark model proposed by Google in 2017 and is also a key technology in the AI ​​language revolution. Previous models were based on recurrent neural networks (RNN, LSTM, etc.). In essence, RNN processes data in a serial manner. Compared with this serial mode, the great innovation of Transformer lies in parallel language processing. At present, some scholars have pioneered the cross-domain application of the Transformer model to computer vision tasks and achieved good results. This is also considered by many AI scholars to have ushered in a new era in the field of CV, and may even completely replace traditional convolution operations.

[0066] Exemplarily, an image processing device needs to identify and locate a user's hand in order to subsequently recognize the actions of the user's hand. Therefore, the target object is usually the user's hand. The shape features, color features, and texture features in the obtained target feature information all belong to the features of the user's hand in the fisheye image, while the background features belong to the features of the non-hand region of the user in the fisheye image. The image processing device processes the target feature information to obtain the target position information of the user's hand in the fisheye image, thereby realizing the identification and location of the user's hand in the fisheye image.

[0067] Specifically, step 103 includes:

[0068] Step 1031, encoding each target feature in the target feature information through the prediction model to obtain encoded feature information;

[0069] In this step, the image processing device encodes each target feature in the target feature information through the prediction model to obtain encoded feature information. Specifically, after obtaining the target feature information corresponding to the fisheye image, for each target feature in the target feature information, the image processing device encodes the target feature through the Encoder (encoder) in the conversion layer of the prediction model to obtain encoded feature information. The target feature information includes multiple downsampled feature maps corresponding to the shape features, color features, texture features, and background features after normalization processing. The specific calculation steps of the Encoder are as follows: for each input downsampled feature map, the Encoder only focuses on the more meaningful positions that the network considers to contain more local information, and uses the fixed number of positions obtained through training as keys to alleviate the problem of large computational complexity caused by large feature maps. In the specific calculation process, the feature map is input to a linear mapping, and 3MK channels are output, where M is the number of detection heads and K is the number of keys. The first 2MK channels encode the sampled offsets, determining which keys the feature map should find. The last MK channels output the contribution (importance index) of each key, and only the contributions of the found keys are normalized. For one feature map, K points are sampled in all layers, fusing the features of different layers to obtain the encoded feature information corresponding to each feature map.

[0070] Step 1032, decoding each encoded feature in the encoded feature information based on the historical position information and the prediction model to obtain decoded feature information;

[0071] In this step, the image processing device decodes each encoded feature in the encoded feature information based on the historical position information and the prediction model to obtain decoded feature information. Specifically, after obtaining the encoded feature information corresponding to each feature map, the image processing device first performs a convolution and a normalization layer on the historical position information of the previous frame of the fisheye image of the current frame to obtain a normalized feature map, then divides the normalized feature map into multiple sub-feature maps, inputs the multiple sub-feature maps into the Decoder in the conversion layer of the prediction model, and at the same time sequentially inputs each encoded feature in the encoded feature information into the Decoder in the conversion layer of the prediction model. For each encoded feature, the multiple sub-feature maps are fused with each encoded feature through the Decoder in the conversion layer of the prediction model to obtain a fused encoded feature, and then a process such as the specific calculation steps of the Encoder is performed on the fused encoded feature to obtain decoded feature information. It can be understood that the decoded feature information includes multiple downsampled feature maps corresponding to shape features, color features, texture features, and background features. After encoding and decoding, the feature expression of the downsampled feature map is more accurate, which helps to improve the accuracy of subsequent recognition and positioning of target objects in the fisheye image.

[0072] Step 1033, classify and regress the decoded feature information through the prediction model to obtain target position information.

[0073] In this step, the image processing device classifies and regresses the decoded feature information through the prediction model to obtain target position information. Specifically, the image processing device processes each downsampled feature map included in the decoded feature information through the conversion layer in the prediction model according to the preset number of bounding boxes. Each downsampled feature map outputs the preset number of bounding box images, then predicts the confidence of each bounding box image of the bounding box images, compares the confidence with the threshold, and only retains the bounding box images with a confidence greater than the threshold as valid bounding box images; for each valid bounding box image, the corresponding Gaussian heat map H is generated as:

[0074]

[0075] where (x,y) are the point coordinates on the image, (x m ,y m ) are the coordinates of the center point of the m-th bounding box, M is the number of valid bounding boxes; the value of sigma is 10, and r represents the radiation range of the center point.

[0076] The image processing device inputs each obtained Gaussian heat map into the fully connected layer and the multi-layer perceptron layer in the prediction model for classification and regression processing to obtain the target position information of the target object.

[0077] Further, step 1033 includes:

[0078] Step 10331, classifying and regressing each decoded feature in the decoded feature information through the prediction model to obtain a prediction confidence level and prediction position information corresponding to each decoded feature;

[0079] In this step, the image processing device inputs each obtained Gaussian heat map into the fully connected layer and multi-layer perceptron layer in the prediction model. For each Gaussian heat map, through the classification head of the fully connected layer, classification processing is performed to predict and output the prediction confidence level of the target object corresponding to each Gaussian heat map in the Gaussian heat map set. Through the regression head of the multi-layer perceptron layer, regression processing is performed to predict and output the prediction position information of the target object corresponding to each Gaussian heat map in the Gaussian heat map set. Among them, the position information of the target object refers to the position coordinates of the target object in the fisheye image.

[0080] Step 10332, determining a target confidence level among the prediction confidence levels corresponding to each decoded feature, and determining the prediction position information corresponding to the target confidence level as the target position information.

[0081] In this step, the image processing device compares the prediction confidence levels of the target objects corresponding to each Gaussian heat map, selects the maximum prediction confidence level as the target confidence level, and determines the prediction position information of the target object corresponding to the Gaussian heat map corresponding to the target confidence level as the target position information of the target object.

[0082] The image processing device in this embodiment obtains a fisheye image and the historical position information of the target object in the fisheye image; performs a first process on the fisheye image through a prediction model to obtain target feature information corresponding to the fisheye image; performs a second process on the target feature information and the historical position information through the prediction model to obtain the target position information of the target object. By predicting the fisheye image through a pre-trained prediction model to obtain the target position information of the target object in the fisheye image, compared with the existing method that requires distorting and correcting the fisheye image, obtaining the corrected image, and then performing recognition and positioning of the target object based on the corrected image, the efficiency of recognizing and positioning the target object based on the fisheye image is improved.

[0083] Reference Figure 2 Combined with the first embodiment, a second embodiment of this image processing method is proposed. The specific process of the second embodiment of this image processing method is the training method of the prediction model. The training method of the prediction model includes the following steps:

[0084] Step a, obtaining training images and a pre-trained prediction network;

[0085] In this step, the image processing device acquires the original fish-eye image, annotates the bounding boxes of the original fish-eye image to obtain the training images, and trains a preliminary lightweight neural network model using the stochastic gradient descent algorithm as the pre-trained prediction network. The training images include a large amount of image data for training the pre-trained prediction network, and the training images are fish-eye images.

[0086] Step b: Input the training images into the pre-trained prediction network to obtain the prediction confidence and prediction position information.

[0087] In this step, the pre-trained prediction network inputs the training images into the pre-trained prediction network to obtain the prediction confidence and prediction position information of the target objects in the pre-trained prediction network.

[0088] Step c: Determine the first loss value based on the prediction confidence and the annotation confidence of the training images, and determine the second loss value based on the prediction position information and the annotation position information of the training images.

[0089] In this step, the image processing device determines the first loss value based on the prediction confidence and the annotation confidence of the training images, and determines the second loss value based on the prediction position information and the annotation position information of the training images.

[0090] Step d: Perform backpropagation based on the first loss value and the second loss value to update the network parameters of the pre-trained prediction network.

[0091] In this step, the image processing device compares the first loss value and the second loss value with the target loss value respectively. If the first loss value or the second loss value is greater than the target loss value, then perform backpropagation based on the first loss value and the second loss value to update the network parameters of the pre-trained prediction network. Among them, the network parameters of the pre-trained prediction network that need to be updated include the relevant parameters in the convolutional layer, pooling layer, normalization layer, transformation layer, fully connected layer, and multi-layer perceptron layer.

[0092] Step e: Until it is determined that the pre-trained prediction network meets the target convergence condition based on the first loss value and the second loss value, then obtain the prediction model.

[0093] In this step, until the image processing device determines that both the first loss value and the second loss value are less than the target loss value, then it is determined that the pre-trained prediction network meets the target convergence condition, and the prediction model is obtained.

[0094] The image processing device of this embodiment trains a pre-trained prediction network to obtain a prediction model, and subsequently uses the pre-trained prediction model to predict a fisheye image to obtain the target position information of the target object in the fisheye image. Compared with the existing method that requires distortion correction of the fisheye image to obtain the corrected image and then performs recognition and positioning of the target object based on the corrected image, the efficiency of recognizing and positioning the target object based on the fisheye image is improved.

[0095] This embodiment also provides an image processing device, which can be specifically integrated in devices such as smart terminals, PC terminals, AR / VR devices, etc. As Figure 3 shown, the image processing device may include:

[0096] An acquisition unit 1001, an acquisition unit, configured to acquire the image information to be processed;

[0097] A first processing unit 1002, configured to perform a first processing on the image information to be processed through the prediction model to obtain target feature information;

[0098] A second processing unit 1003, configured to perform a second processing on the target feature information to obtain target position information.

[0099] In an optional example, the first processing unit is further configured to:

[0100] Perform convolution processing and pooling processing on the image information to be processed through the prediction model to obtain first feature information;

[0101] Perform normalization processing on the first feature information through the prediction model to obtain target feature information.

[0102] In an optional example, the first processing unit is further configured to:

[0103] Perform convolution processing on the image information to be processed through the prediction model to obtain semantic feature information;

[0104] Perform pooling processing on each semantic feature in the semantic feature information through the prediction model to obtain first feature information.

[0105] In an optional example, the first processing unit is further configured to:

[0106] Calculate the mean and variance corresponding to the first feature information through the prediction model;

[0107] Based on the mean, the variance, and the target weight, perform normalization processing on each first feature in the first feature information through the prediction model to obtain target feature information.

[0108] In an optional example, the second processing unit is further configured to:

[0109] Encode each target feature in the target feature information through the prediction model to obtain encoded feature information;

[0110] Decode each encoded feature in the encoded feature information based on the historical position information and the prediction model to obtain decoded feature information;

[0111] Classify and regress the decoded feature information through the prediction model to obtain target position information.

[0112] In an optional example, the second processing unit is further configured to:

[0113] Classify and regress each decoded feature in the decoded feature information through the prediction model to obtain a prediction confidence level and prediction position information corresponding to each decoded feature;

[0114] Determine a target confidence level among the prediction confidence levels corresponding to each decoded feature, and determine the prediction position information corresponding to the target confidence level as the target position information.

[0115] In an optional example, the image processing device further includes a training unit, and the training unit is configured to:

[0116] Obtain a training image and a pre-trained prediction network;

[0117] Input the training image into the pre-trained prediction network to obtain prediction confidence level and prediction position information;

[0118] Determine a first loss value based on the prediction confidence level and the annotation confidence level of the training image, and determine a second loss value based on the prediction position information and the annotation position information of the training image;

[0119] Perform backpropagation based on the first loss value and the second loss value to update the network parameters of the pre-trained prediction network;

[0120] Until it is determined that the pre-trained prediction network meets the target convergence condition based on the first loss value and the second loss value, then obtain the prediction model.

[0121] Adopt the solution of this embodiment to obtain the image information to be processed; perform a first process on the image information to be processed through a prediction model to obtain target feature information; perform a second process on the target feature information to obtain target position information. The image information to be processed is a fisheye image. By using a pre-trained prediction model to predict the fisheye image, the target position information of the target object in the fisheye image is obtained. Compared with the prior art that requires distortion correction of the fisheye image to obtain the corrected image and then perform recognition and positioning of the target object based on the corrected image, the efficiency of recognizing and positioning the target object based on the fisheye image is improved.

[0122] Correspondingly, an embodiment of the present application further provides an image processing device, as Figure 4 shown, Figure 4 is a schematic structural diagram of the image processing device provided by the embodiment of the present application. The image processing device 1100 includes a processor 1101 with one or more processing cores, a memory 1102 with one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structure of the image processing device shown in the figure does not constitute a limitation on the image processing device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0123] The processor 1101 is the control center of the image processing device 1100. It uses various interfaces and lines to connect all parts of the entire image processing device 1100. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, it executes various functions of the image processing device 1100 and processes data, thereby monitoring the entire image processing device 1100. The processor 1101 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logical block diagrams disclosed in the embodiments of the present application.

[0124] In the embodiment of the present application, the processor 1101 in the image processing device 1100 will load the computer programs corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to execute any one of the image processing methods provided by the embodiment of the present application. For the specific implementation of each operation of the image processing method, reference can be made to the previous embodiments and will not be elaborated here.

[0125] Optionally, as Figure 4As shown, the image processing device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is respectively electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107. Those skilled in the art can understand that Figure 4 the structure of the image processing device shown in does not constitute a limitation on the image processing device, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0126] The touch display screen 1103 can be used to display a graphical user interface and receive operation computer programs generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the image processing device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation computer programs, and the operation computer programs execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1101, and can receive and execute the command sent by the processor 1101. The touch panel can cover the display panel. After the touch panel detects a touch operation on or near it, it transmits it to the processor 1101 to determine the type of touch event. Subsequently, the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to implement the input function.

[0127] The radio frequency circuit 1104 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other image processing devices, and to receive and transmit signals between the network device or other image processing devices.

[0128] The audio circuit 1105 can be used to provide an audio interface between the user and the image processing device through a speaker and a microphone. The audio circuit 1105 can convert the received electrical signal after the audio data is converted, and transmit it to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and then converted into audio data. After the audio data is output to the processor 1101 for processing, it is transmitted by the radio frequency circuit 1104 to, for example, another image processing device, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the image processing device.

[0129] The input unit 1106 can be used to receive input digital, character information or user target feature information (such as fingerprint, iris, face information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0130] The power supply 1107 is used to supply power to each component of the image processing device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management device, so as to realize functions such as management of charging, discharging, and power consumption management through the power management device. The power supply 1107 may also include any components such as one or more DC or AC power supplies, recharge devices, power failure detection circuits, power converters or inverters, and power status indicators.

[0131] Although Figure 4 not shown in the figure, the image processing device 1100 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.

[0132] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by a computer program, or by controlling relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0134] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple computer programs that can be loaded by a processor to execute any one of the image processing methods provided by the embodiments of the present application. For the specific implementation of each operation of the image processing method, reference can be made to the previous embodiments and will not be elaborated here.

[0135] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0136] Since the computer programs stored in the computer-readable storage medium can execute any one of the image processing methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the image processing methods provided by the embodiments of the present application can be realized. For details, reference can be made to the previous embodiments and will not be elaborated here.

[0137] According to one aspect of the present application, there is also provided a computer program product or a computer program. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the image processing device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the image processing device executes the methods provided in the various optional implementation manners in the above embodiments.

[0138] In the above embodiments of the image processing device, computer-readable storage medium, image processing device, and computer program product, the descriptions of each embodiment have their own focuses. For parts not elaborated in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects brought by the above-described image processing device, computer-readable storage medium, computer program product, image processing device, and their corresponding units can refer to the description of the image processing method in the above embodiments and will not be elaborated here specifically.

[0139] The above has introduced in detail an image processing method, device, equipment, computer-readable storage medium, and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An image processing method, characterized in that, The described image processing method includes: Obtain the image information to be processed; Perform a first processing on the image information to be processed through a prediction model to obtain target feature information; Perform a second processing on the target feature information to obtain target position information.

2. The image processing method according to claim 1, wherein The step of performing a first processing on the image information to be processed through a prediction model to obtain target feature information includes: Perform convolution processing and pooling processing on the image information to be processed through a prediction model to obtain first feature information; Perform normalization processing on the first feature information through a prediction model to obtain target feature information.

3. The image processing method according to claim 2, wherein The step of performing convolution processing and pooling processing on the image information to be processed through a prediction model to obtain first feature information includes: Perform convolution processing on the image information to be processed through a prediction model to obtain semantic feature information; Perform pooling processing on each semantic feature in the semantic feature information through a prediction model to obtain first feature information.

4. The image processing method according to claim 2, wherein The step of performing normalization processing on the first feature information through a prediction model to obtain target feature information includes: Calculate the mean and variance corresponding to the first feature information through a prediction model; Based on the mean, the variance, and a target weight, perform normalization processing on each first feature in the first feature information through a prediction model to obtain target feature information.

5. The image processing method according to claim 1, characterized in that The step of performing a second processing on the target feature information to obtain target position information includes: Perform encoding processing on each target feature in the target feature information through the prediction model to obtain encoded feature information; Based on historical position information and the prediction model, perform decoding processing on each encoded feature in the encoded feature information to obtain decoded feature information; Perform classification processing and regression processing on the decoded feature information through the prediction model to obtain target position information.

6. The image processing method according to claim 5, wherein The step of performing classification processing and regression processing on the decoded feature information through the prediction model to obtain target position information includes: Perform classification processing and regression processing on each decoded feature in the decoded feature information through the prediction model to obtain the predicted confidence and predicted position information corresponding to each decoded feature; Determine a target confidence in the predicted confidences corresponding to each decoded feature, and determine the predicted position information corresponding to the target confidence as the target position information.

7. The image processing method according to any one of claims 1-6, characterized in that, The training method of the prediction model includes: Obtain training images and a pre-trained prediction network; Input the training images into the pre-trained prediction network to obtain predicted confidence and predicted position information; Determine a first loss value based on the predicted confidence and the labeled confidence of the training images, and determine a second loss value based on the predicted position information and the labeled position information of the training images; Perform backpropagation based on the first loss value and the second loss value to update the network parameters of the pre-trained prediction network; Until it is determined that the pre-trained prediction network meets the target convergence condition based on the first loss value and the second loss value, then obtain the prediction model.

8. An image processing apparatus, characterized in that, The device includes: An acquisition unit for obtaining the image information to be processed; The first processing unit is configured to perform a first processing on the to-be-processed image information through a prediction model to obtain target feature information; The second processing unit is configured to perform a second processing on the target feature information to obtain target position information; Preferably, the first processing unit performing a first processing on the to-be-processed image information through a prediction model to obtain target feature information includes: Performing a convolution processing and a pooling processing on the to-be-processed image information through the prediction model to obtain first feature information; Performing a normalization processing on the first feature information through the prediction model to obtain target feature information; Preferably, the first processing unit performing a convolution processing and a pooling processing on the to-be-processed image information through the prediction model to obtain first feature information includes: Performing a convolution processing on the to-be-processed image information through the prediction model to obtain semantic feature information; Performing a pooling processing on each semantic feature in the semantic feature information through the prediction model to obtain first feature information; Preferably, the first processing unit performing a normalization processing on the first feature information through the prediction model to obtain target feature information includes: Calculating a mean value and a variance corresponding to the first feature information through the prediction model; Based on the mean value, the variance and a target weight, performing a normalization processing on each first feature in the first feature information through the prediction model to obtain target feature information; Preferably, the second processing unit performing a second processing on the target feature information to obtain target position information includes: Performing an encoding processing on each target feature in the target feature information through the prediction model to obtain encoded feature information; Based on historical position information and the prediction model, performing a decoding processing on each encoded feature in the encoded feature information to obtain decoded feature information; Performing a classification processing and a regression processing on the decoded feature information through the prediction model to obtain target position information; Preferably, the second processing unit performing a classification processing and a regression processing on the decoded feature information through the prediction model to obtain target position information includes: Performing a classification processing and a regression processing on each decoded feature in the decoded feature information through the prediction model to obtain a prediction confidence and prediction position information corresponding to each decoded feature; Determining a target confidence in the prediction confidences corresponding to each decoded feature, and determining the prediction position information corresponding to the target confidence as the target position information; Preferably, the training method of the prediction model by the first processing unit includes: Obtaining training images and a pre-trained prediction network; Inputting the training images into the pre-trained prediction network to obtain prediction confidence and prediction position information; Determining a first loss value based on the prediction confidence and the annotation confidence of the training images, and determining a second loss value based on the prediction position information and the annotation position information of the training images; Performing backpropagation based on the first loss value and the second loss value to update network parameters of the pre-trained prediction network; Until it is determined based on the first loss value and the second loss value that the pre-trained prediction network meets the target convergence condition, the prediction model is obtained.

9. An image processing apparatus, characterized in that, The invention comprises a processor and a memory, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor executes the steps of any one of the methods of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a computer program, and when the computer program is run on an image processing device, the computer program is used to cause the image processing device to perform the steps of the method according to any one of claims 1 to 7.