Image Processing Method, Apparatus, Device, and Storage Medium
By extracting condition parameter information from consecutive frames and performing cross-mask recognition training, the method addresses the issue of low accuracy in video sequence image processing, achieving improved segmentation and recognition.
Patent Information
- Application Number
- CN202110348195.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-31
AI Technical Summary
Prior art In video sequence image segmentation, single frame image processing results in missing image reference details of other frames, reducing the accuracy of image processing.
By obtaining the continuous target frame image and adjacent frame image under the video sequence sample, the conditional parameter information of the instance information is extracted, and cross-mask recognition training is performed, and the trained preset model is obtained, which is used to mask recognition processing for the video sequence to be recognized.
The accuracy of image processing is improved, and the accuracy of image segmentation is significantly improved by learning image information across frames.
Smart Images

Figure CN113705307B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] With the continuous development of computer technologies, image processing technologies based on artificial intelligence have become increasingly mature. Image segmentation is a crucial preprocessing step for image recognition and computer vision, and is widely applied in various fields. For example, it can effectively assist tasks such as image classification, object detection, and object tracking in various scene images.
[0003] In the prior art, image detection and segmentation are usually performed by means of morphological matching or template matching. However, for video sequences, performing image segmentation on a single-frame image will result in the loss of reference details on the images of other frames during the segmentation process, leading to a low accuracy of image processing. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, apparatus, device, and storage medium, which can improve the accuracy of image processing.
[0005] To solve the above technical problems, the embodiments of this application provide the following technical solutions:
[0006] An image processing method, comprising:
[0007] Obtaining consecutive target frame images and adjacent frame images in a video sequence sample;
[0008] Extracting first conditional parameter information corresponding to instance information in the target frame image and second conditional parameter information corresponding to instance information in the adjacent frame image;
[0009] Inputting the first conditional parameter information and the second conditional parameter information into a preset model, and performing cross-mask recognition training on the target frame image and the adjacent frame image to obtain a trained preset model;
[0010] Performing mask recognition processing on a video sequence to be recognized based on the trained preset model.
[0011] An image processing apparatus, comprising:
[0012] An obtaining unit, configured to obtain consecutive target frame images and adjacent frame images in a video sequence sample;
[0013] An extracting unit, configured to extract first conditional parameter information corresponding to instance information in the target frame image and second conditional parameter information corresponding to instance information in the adjacent frame image;
[0014] An input unit, configured to input the first conditional parameter information and the second conditional parameter information into a preset model, perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtain a trained preset model;
[0015] A recognition unit, configured to perform mask recognition processing on the video sequence to be recognized based on the trained preset model.
[0016] In some embodiments, the obtaining unit is configured to:
[0017] Obtain a target frame image and an adjacent frame image with an interval time not exceeding a preset time threshold in a video sequence sample.
[0018] In some embodiments, the recognition unit is configured to:
[0019] Obtain an image to be recognized for each frame in the video sequence to be recognized;
[0020] Extract third conditional parameter information corresponding to instance information in the image to be recognized;
[0021] Input the third conditional parameter information into the trained preset model, and output corresponding target mask information.
[0022] In some embodiments, the image processing device further includes:
[0023] A sample recognition unit, configured to recognize sample mask information corresponding to instance information in a sample image for each frame in a video sequence sample based on the trained preset model;
[0024] A classification training unit, configured to input the sample mask information into a preset classification model for classification training to obtain a trained preset classification model.
[0025] In some embodiments, the classification training unit is configured to:
[0026] Input the sample mask information into a preset classification model;
[0027] Convert the sample mask information into a vector representation space of a corresponding preset dimension;
[0028] Classify the instance information in the sample image for each frame according to the similarity between the vector representation spaces;
[0029] Iteratively adjust the classification network parameters in the preset classification model according to a second difference value between the classification result and the label classification result until the second difference value converges, and obtain a trained preset classification model.
[0030] In some embodiments, the image processing device further includes:
[0031] A classification unit for classifying instance information in a video sequence to be recognized according to the trained preset classification model, so as to implement instance segmentation of the video sequence to be recognized.
[0032] A computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the above image processing method.
[0033] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above image processing method are implemented.
[0034] A computer program product or a computer program includes computer instructions stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer implements the steps in the above image processing method.
[0035] In the embodiments of the present application, by obtaining an image to be processed, and by obtaining consecutive target frame images and adjacent frame images in a video sequence sample; extracting first conditional parameter information corresponding to instance information in the target frame image and second conditional parameter information corresponding to instance information in the adjacent frame image; inputting the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtaining a trained preset model; performing mask recognition processing on the video sequence to be recognized based on the trained preset model. In this way, adjacent target frame images and adjacent frame images in the video sequence sample can be obtained, and the corresponding conditional parameter information of both can be extracted as convolution kernels respectively to perform cross-mask recognition learning, and a trained preset model with more accurate mask recognition is obtained for recognition. Compared with the solution of performing image recognition and segmentation on a single frame image, the embodiments of the present application can learn cross-frame image information, thereby greatly improving the accuracy of image processing. Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0037] Figure 1 It is a schematic diagram of the scenario of the image processing system provided by the embodiments of the present application;
[0038] Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0039] Figure 3 is another schematic flowchart of the image processing method provided by an embodiment of the present application;
[0040] Figure 4 is a schematic diagram of a scenario of the image processing method provided by an embodiment of the present application;
[0041] Figure 5 is a schematic structural diagram of an image processing apparatus provided by an embodiment of the present application;
[0042] Figure 6 is a schematic structural diagram of a server provided by an embodiment of the present application. Specific Embodiments
[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0044] An embodiment of the present application provides an image processing method, apparatus, device, and storage medium.
[0045] Please refer to Figure 1 , Figure 1 is a schematic diagram of a scenario of an image processing system provided by an embodiment of the present application, including: terminal A and a server (the image processing system may further include other terminals other than terminal A, and the specific number of terminals is not limited here). Terminal A and the server can be connected through a communication network, and the communication network can include a wireless network and a wired network, where the wireless network includes one or a combination of a wireless wide area network, a wireless local area network, a wireless metropolitan area network, and a wireless personal area network. Network entities such as routers and gateways are included in the network, which are not shown in the figure. Terminal A can interact with the server through the communication network. For example, terminal A can send a video sequence to be recognized to the server.
[0046] The image processing system may include an image processing apparatus, and the image processing apparatus may be specifically integrated in the server. In some embodiments, the image processing apparatus may also be integrated in a terminal with computing capabilities. In this embodiment, it is described that the image processing apparatus is integrated in the server, as Figure 1As shown, the server obtains consecutive target frame images and adjacent frame images in the video sequence sample; extracts first conditional parameter information corresponding to the instance information in the target frame image and second conditional parameter information corresponding to the instance information in the adjacent frame image; inputs the first conditional parameter information and the second conditional parameter information into a preset model, performs cross-mask recognition training on the target frame image and the adjacent frame image, and obtains a trained preset model; performs mask recognition processing on the video sequence to be recognized sent by the receiving terminal A based on the trained preset model.
[0047] The image processing system may further include a terminal A, which can install various applications required by users, such as instant video processing applications, media applications, and browser applications. The terminal A can upload the video sequence to be recognized to the server for review.
[0048] It should be noted that Figure 1 The scene schematic diagram of the image processing system shown is only an example. The image processing system and scene described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the image processing system and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0049] The following will be described in detail respectively.
[0050] In this embodiment, it will be described from the perspective of an image processing device, which can be specifically integrated in a server equipped with a storage unit and installed with a microprocessor and having computing power.
[0051] Please refer to Figure 2 , Figure 2 is a flowchart of the image processing method provided by the embodiments of the present application. The image processing method includes:
[0052] In step 101, consecutive target frame images and adjacent frame images in the video sequence sample are obtained.
[0053] Among them, the video sequence sample is video information prepared in advance for training. The video information is composed of multiple consecutive frames of images. Generally, the playback rate of the video information is 24 frames per second to form a continuous picture.
[0054] It can be understood that video instance segmentation generally refers to segmenting each instance information in the video frames of a video. This instance information can refer to object information in the video frames, such as object information like human information, animal information, etc. However, in current video instance segmentation research, single-frame video frames are mainly used for detection and segmentation, often ignoring the instance details of other frames inherent in the video, making video instance segmentation often inaccurate.
[0055] In order to solve the above technical problems in the embodiments of the present application, the target frame image and the adjacent frame image that are consecutive in frame playback under the video sequence sample can be obtained. Since the target frame image and the adjacent frame image are consecutive in time, and the instance information on the target frame image and the adjacent frame image is consecutive in playback, the instance information of both has a large number of identical features. At the same time, because the target frame image and the adjacent frame image are different in playback, the instance information of both also has subtle feature differences, and this subtle feature difference can be used as the learning direction for the subsequent model.
[0056] In some embodiments, the obtaining of the consecutive target frame image and adjacent frame image under the video sequence sample may include: obtaining the target frame image and the adjacent frame image with an interval time not exceeding a preset time threshold under the video sequence sample.
[0057] Among them, in order to increase the flexibility of image selection, the preset time threshold can be set. This preset time threshold is the critical value for defining whether the target frame image and the adjacent frame image are consecutive, and it can be set by the user. For example, the time corresponding to playing 20 frames of images. In this way, the embodiments of the present application can obtain the target frame image and the adjacent frame image with an interval time not exceeding the preset time threshold under the video sequence sample.
[0058] In step 102, the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image are extracted.
[0059] Among them, the first conditional parameter information can be understood as the convolution kernel information for extracting the features corresponding to the instance information in the target frame image, and the second conditional parameter information can be understood as the convolution kernel information for extracting the features corresponding to the instance information in the adjacent frame image. The convolution kernel is the filter.
[0060] In one embodiment, the target frame image and the adjacent frame images can be pre-processed by fully convolutional network to obtain corresponding first image features and second image features. The fully convolutional network (FCN) processing refers to pixel-level classification of images (i.e., each pixel is classified), thus solving the problem of semantic-level image segmentation. The fully convolutional network processing can accept input images of any size, and uses a transposed convolution layer to upsample the feature map of the last convolutional layer to restore it to the same size as the input image, so as to generate a prediction for each pixel while retaining the spatial information in the original input image.
[0061] Thereby, convolutional processing can be performed on the first image features to extract first conditional parameters regarding instance information in the target frame image, and convolutional processing can be performed on the second image features to extract second conditional parameter information regarding instance information in the target frame image. The convolutional processing can be implemented by a Convolutional Neural Networks (CNN).
[0062] In step 103, the first conditional parameter information and the second conditional parameter information are input into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame images, thereby obtaining a trained preset model.
[0063] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0064] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0065] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision research focuses on related theories and technologies, aiming to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0066] Among them, in the embodiments of the present application, corresponding image semantic processing is combined with computer vision technology. The mask can be the bitmap position occupied by the instance information in the image on the image, and instance matting of the instance information can be performed through this mask. The preset model can be a feature extraction network model, such as the DLA-34 network model, etc. Thus, the first conditional parameter information and the second conditional parameter information are input into the preset model. Using the first conditional parameter as the convolution kernel, learning of mask recognition is respectively performed on the instance information in the target frame image and the adjacent frame image, and using the second conditional parameter as the convolution kernel, learning of mask recognition is respectively performed on the instance information in the adjacent frame image and the target frame image, so that the preset model can learn the features of the instance information in different pose states, that is, learn the mask information of the instance information in different positions, until the learning expression of the mask recognition converges, and a trained preset model is obtained.
[0067] In some embodiments, the step of inputting the first conditional parameter information and the second conditional parameter information into the preset model and performing cross-mask recognition training on the target frame image and the adjacent frame image to obtain a trained preset model may include:
[0068] (1) Performing mask convolution processing on the first image feature to obtain first mask feature information;
[0069] (2) Performing mask convolution processing on the second image feature to obtain second mask feature information;
[0070] (3) Inputting the first conditional parameter information and the second conditional parameter information into the preset model, and performing cross-mask recognition training on the first mask feature information and the second mask feature information to obtain a trained preset model.
[0071] Among them, in order to make the subsequent mask recognition process more accurate, the first image feature can be pre-processed by mask convolution. This mask convolution process can be processed by a first preset convolutional network, which can be composed of four stacked 3x3 convolutional layers. After this convolutional process, a first mask feature biased towards the semantic expression of the mask can be obtained. Similarly, by performing mask convolution processing on the second image feature, a second mask feature biased towards the semantic expression of the mask can be obtained.
[0072] Furthermore, using the first conditional parameter as the convolutional kernel, learning mask recognition is respectively performed on the instance information in the first mask feature and the second mask feature, and using the second conditional parameter as the convolutional kernel, learning mask recognition is respectively performed on the instance information in the adjacent frame image and the target frame image, so that the preset model can better learn the features of the instance information in different pose states, that is, learn the mask information of the instance information at different positions, until the learning expression of the mask recognition converges, and a trained preset model is obtained.
[0073] In some embodiments, the training process, that is, performing cross-mask recognition training on the first mask feature information and the second mask feature information to obtain a trained preset model, may include:
[0074] (1.1) Performing convolutional processing on the first mask feature information and the second mask feature information respectively through the first conditional parameter information to obtain first mask information and first cross-mask information;
[0075] (1.2) Performing convolutional processing on the first mask feature information and the second mask feature information respectively through the second conditional parameter information to obtain second mask information and second cross-mask information;
[0076] (1.3) Calculating a first difference value between the first mask information, the first cross-mask information, the second mask information, and the second cross-mask information and the corresponding label information;
[0077] (1.4) Iteratively adjusting the network model parameters of the preset model according to the first difference value until the first difference value converges, and obtaining a trained preset model.
[0078] Among them, the first conditional parameter can be used as the convolutional kernel to perform multiple convolutional processes on the first mask feature information and the second mask feature information respectively to obtain the first mask information and the first cross-mask information corresponding to the instance information in the target frame image and the adjacent frame image. Similarly, the second conditional parameter can be used as the convolutional kernel to perform multiple convolutional processes on the second mask feature information and the first mask feature information respectively to obtain the second mask information and the second cross-mask information corresponding to the instance information in the adjacent frame image and the target frame image.
[0079] Further, the tag information is the true mask information of the instance information in the target frame image and the adjacent frame images. The true mask information can be manually calibrated. Thus, the first difference values between the first mask information, the first cross-mask information, the second mask information, and the second cross-mask information and the corresponding tag information can be calculated, and the network model parameters of the preset model can be iteratively adjusted until the first difference value converges, obtaining the trained preset model. In the embodiments of the present application, by cross-learning the continuous feature properties between video frames, the trained preset model can learn the mask information in different pose states. Compared with the solution of mask recognition for a single-frame image, the trained preset model can greatly improve the accuracy of mask information recognition, and thus greatly improve the accuracy of image segmentation processing.
[0080] In step 104, mask recognition processing is performed on the video sequence to be recognized based on the trained preset model.
[0081] Since the accuracy of the trained preset model in recognizing the mask information of the instance information in the image is much higher than that of mask recognition for a single-frame image, the instance information in each frame of the video sequence to be recognized can be recognized based on the trained preset model, obtaining all the instance information of the video sequence to be recognized, so as to implement subsequent instance segmentation of the video sequence to be recognized.
[0082] In some embodiments, the mask recognition processing of the video sequence to be recognized based on the trained preset model may include:
[0083] (1) Obtain the image to be recognized in each frame of the video sequence to be recognized;
[0084] (2) Extract the third conditional parameter information corresponding to the instance information in the image to be recognized;
[0085] (3) Input the third conditional parameter information into the trained preset model and output the corresponding target mask information.
[0086] The image to be recognized in each frame of the video sequence to be recognized can be obtained in advance, and the third conditional parameter information corresponding to the instance information in each image to be recognized can be extracted according to the above-mentioned conditional parameter information extraction method. Thus, the third conditional parameter information is input into the trained preset model in sequence, and the target mask information corresponding to the instance information in each image to be recognized is output, so as to implement subsequent instance segmentation of the video sequence to be recognized.
[0087] As can be seen from the above, in the embodiment of the present application, by obtaining the image to be processed, and obtaining consecutive target frame images and adjacent frame images in the video sequence sample; extracting the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image; inputting the first conditional parameter information and the second conditional parameter information into a preset model, performing cross-mask recognition training on the target frame image and the adjacent frame image, and obtaining the trained preset model; performing mask recognition processing on the video sequence to be recognized based on the trained preset model. In this way, adjacent target frame images and adjacent frame images in the video sequence sample can be obtained, and the corresponding conditional parameter information of the two can be respectively extracted as convolution kernels to perform cross-mask recognition learning on them, and a trained preset model with more accurate mask recognition can be obtained for recognition. Compared with the solution of performing image recognition and segmentation on a single frame image, the embodiment of the present application can learn cross-frame image information, thereby greatly improving the accuracy of image processing.
[0088] The following will give further detailed descriptions by way of examples.
[0089] In this embodiment, it will be described by taking the specific integration of the image processing device in the server as an example. The embodiment of the present application will be described by taking the identification scenario of product identification as an example, and specific references are as follows.
[0090] Please refer to Figure 3 , Figure 3 which is another schematic flowchart of the image processing method provided by the embodiment of the present application. The method process may include:
[0091] In step 201, the server obtains a target frame image and an adjacent frame image in the video sequence sample with an interval time not exceeding a preset time threshold.
[0092] Among them, in order to better illustrate the embodiment of the present application, please refer to Figure 4 shown in Figure 4 which is a schematic diagram of the scenario of the image processing method provided by the embodiment of the present application. The preset time threshold is a critical value for defining whether the target frame image and the adjacent frame image are consecutive, which can be 0.8 seconds. The server can randomly obtain the target frame image 11 and the adjacent frame image 12 in the video sequence sample with an interval time not exceeding 0.8 seconds. It can be seen that the instance information in the target frame image 11 and the adjacent frame image 12 is both human, but the instance information of the two has different pose states.
[0093] In step 202, the server performs full convolution processing on the target frame image and the adjacent frame image to obtain corresponding first image features and second image features.
[0094] Among them, in order to better perform subsequent feature extraction on the target frame image and adjacent frame images, the target frame image and adjacent frame images can be subjected to full convolution processing through a fully convolutional network. The fully convolutional networks of both share the same weights, and first image features and second image features that generate certain predictions for each pixel are obtained.
[0095] In step 203, the server performs convolution processing on the first image features through a convolution layer of a preset size to obtain first conditional parameter information corresponding to the instance information in the first image features, and performs convolution processing on the second image features through a convolution layer of a preset size to obtain second conditional parameter information corresponding to the instance information in the second image features.
[0096] Among them, the preset size can be 1×1 size. The 1×1 size convolution layer can be added with an activation layer to enhance the expression ability by adding non-linear activation. In this way, the server can perform convolution processing on the first image features through a convolution layer of a preset size (i.e., the control head in the figure) to obtain first conditional parameter information θ x,y (t) corresponding to the features expressing instance information in the first image features. Similarly, by performing convolution processing on the second image features through a convolution layer of a preset size (i.e., the control head in the figure), second conditional parameter information θ x′,y′ (t + δ) corresponding to the features expressing instance information in the second image features is obtained. The first conditional parameter information can be understood as the convolution kernel information for extracting the features corresponding to the instance information in the target frame image, and the second conditional parameter information can be understood as the convolution kernel information for extracting the features corresponding to the instance information in the adjacent frame image.
[0097] In step 204, the server performs masked convolution processing on the first image features to obtain first masked feature information, performs masked convolution processing on the second image features to obtain second masked feature information, and inputs the first conditional parameter information and second conditional parameter information into a preset model.
[0098] Among them, in order to make the subsequent masked recognition process more accurate, the first image features can be pre-processed by masked convolution. The masked convolution processing can be processed by a first preset convolutional network (i.e., the masked feature branch in the figure). The first preset convolutional network can be composed of four 3×3 convolution layers stacked. After this convolution processing, first masked features biased towards the semantic expression of the mask can be obtained. Similarly, by performing masked convolution processing on the second image features, second masked features biased towards the semantic expression of the mask can be obtained.
[0099] Furthermore, the first conditional parameter information θ x,y (t) and the second conditional parameter information θx′,y′ The input value of (t + δ) is prepared for training in the preset model.
[0100] In step 205, the server performs convolution processing on the first masked feature information and the second masked feature information respectively through the first conditional parameter information to obtain the first masked information and the first cross-masked information, and performs convolution processing on the first masked feature information and the second masked feature information respectively through the second conditional parameter information to obtain the second masked information and the second cross-masked information.
[0101] Among them, the server can use the first conditional parameter information θ x,y (t) as the convolution kernel, and perform multiple convolution processes on the first masked feature information and the second masked feature information respectively to obtain the first masked information M x,y (t) corresponding to the instance information in the target frame image and the adjacent frame image and the first cross-masked information Similarly, the server can use the second conditional parameter information θ x′,y′ (t + δ) as the convolution kernel, and perform multiple convolution processes on the second masked feature information and the first masked feature information respectively to obtain the second masked information M x′,y′ (t + δ) corresponding to the instance information in the adjacent frame image and the target frame image and the second cross-masked information The specific calculation process can be shown as the following formulas (1), (2), (3), and (4):
[0102]
[0103]
[0104]
[0105]
[0106] The * represents the convolution operation.
[0107] In step 206, the server calculates the first difference value between the first masked information, the first cross-masked information, the second masked information, and the second cross-masked information and the corresponding label information, and iteratively adjusts the network model parameters of the preset model according to the first difference value until the first difference value converges to obtain the trained preset model.
[0108] Among them, the label information is the true masked information of the instance information in the target frame image and the adjacent frame image, and the true masked information can be manually calibrated. Thus, the first masked information M x,y(t), the first cross-mask information The second mask information M x′,y′ (t + δ) and the second cross-mask information The first difference value between the corresponding label information is used to iteratively adjust the network model parameters of the preset model until the first difference value converges, and the trained preset model is obtained. That is, in the embodiments of the present application, by cross-learning the continuous feature properties between video frames, the trained preset model can learn the mask information in different pose states. Compared with the scheme of mask recognition for a single-frame image, the trained preset model can greatly improve the accuracy of mask information recognition, and further greatly improve the accuracy of image segmentation processing.
[0109] In step 207, the server obtains the to-be-recognized image of each frame in the to-be-recognized video sequence, extracts the third conditional parameter information corresponding to the instance information in the to-be-recognized image, and inputs the third conditional parameter information into the trained preset model to output the corresponding target mask information.
[0110] Among them, since the accuracy of mask information recognition of the instance information in the image by the trained preset model is much higher than that of mask recognition for a single-frame image, the third conditional parameter information corresponding to the instance information in each to-be-recognized image is extracted according to the above conditional parameter information extraction method. Therefore, the third conditional parameter information is sequentially input into the trained preset model to output the target mask information corresponding to the instance information in each to-be-recognized image, so as to realize subsequent instance segmentation of the to-be-recognized video sequence.
[0111] In step 208, the server identifies the sample mask information corresponding to the instance information in the sample image of each frame in the video sequence sample based on the trained preset model, and inputs the sample mask information into the preset classification model.
[0112] Among them, in order to realize video instance segmentation, the server can identify the sample mask information corresponding to the instance information in the sample image of each frame in the video sequence sample based on the trained preset model, that is, calibrate the sample mask information corresponding to the instance information in the video sequence sample.
[0113] In order to realize subsequent video instance segmentation and object tracking, the mask information can be input into the preset classification model for non-linear processing, and the preset classification model can be a convolutional neural network model.
[0114] In step 209, the server converts the sample mask information into a corresponding vector representation space of a preset dimension, classifies the instance information in each frame of the sample image according to the similarity between the vector representation spaces, and iteratively adjusts the classification network parameters in the preset classification model according to the second difference value between the classification result and the label classification result until the second difference value converges, obtaining the trained preset classification model.
[0115] Among them, the preset dimension is set and can be 200 dimensions, etc. The server can uniformly express the mask information as a vector representation space of preset dimension N through convolution processing by the preset classification model, and classify the instance information in each frame of the sample image according to the cosine similarity between the N-dimensional vector representation spaces. This cosine similarity, also known as cosine similarity, evaluates their similarity by calculating the cosine value of the included angle between two vectors.
[0116] Thus, the label classification result is the classification result of artificially pre-classifying the same type of mask information into the same category in advance, and this label classification result can be used as a standard for reference. Mask information with a cosine similarity less than the preset cosine similarity can be classified into one category. The preset cosine similarity is the critical value for determining whether the vector representation spaces belong to the same category. For example, 0.1. According to the second difference degree between the classification result of the preset classification model and the label classification result, the classification network parameters in the preset classification model are iteratively adjusted until the second difference value converges, obtaining the trained preset classification model. This trained preset classification model can accurately obtain the N-dimensional vector representation space from the mask information.
[0117] In step 210, the server classifies the instance information in the video sequence to be recognized according to the trained preset classification model.
[0118] Among them, the server inputs the target mask information corresponding to the instance information in each image to be recognized into the trained preset classification model, converts each target mask information into an N-dimensional vector representation space, and classifies the instance information according to the cosine similarity between the N-dimensional vector representation spaces to determine the same instances in different frame images for video instance segmentation and video tracking and other processing. In one embodiment, the following formula can be referred to for classification:
[0119]
[0120] Among them, the p i (n) is one-hot encoding, also known as one-hot encoding. It mainly uses a bit status register to encode each state. Each state has its own independent register bit, and only one bit is valid at any time. One-hot encoding uses 0 and 1 to represent some parameters. The e iRefers to the vector representation space of instance information, where T is the transpose matrix and w j Refers to the weights of the classifier. Here, n refers to the number of instance information, and exp() represents the exponential function with the natural constant e as the base.
[0121] As can be seen from the above, in the embodiment of the present application, by obtaining the image to be processed, and by obtaining consecutive target frame images and adjacent frame images in the video sequence sample; extracting the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image; inputting the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, obtaining the trained preset model; and performing mask recognition processing on the video sequence to be recognized based on the trained preset model. In this way, adjacent target frame images and adjacent frame images in the video sequence sample can be obtained, and the corresponding conditional parameter information of both can be extracted as convolution kernels respectively to perform cross-mask recognition learning on them, obtaining a trained preset model with more accurate mask recognition for recognition. Compared with the scheme of performing image recognition and segmentation on a single frame image, the embodiment of the present application can learn cross-frame image information, thereby greatly improving the accuracy of image processing.
[0122] To facilitate better implementation of the image processing method provided by the embodiment of the present application, the embodiment of the present application also provides a device based on the above image processing method. The meanings of the terms are the same as those in the above image processing method, and the specific implementation details can refer to the description in the method embodiment.
[0123] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of the image processing device provided by the embodiment of the present application. The image processing device may include an acquisition unit 301, an extraction unit 302, an input unit 303, a recognition unit 304, etc.
[0124] The acquisition unit 301 is used to acquire consecutive target frame images and adjacent frame images in the video sequence sample.
[0125] In some embodiments, the acquisition unit 301 is used to: acquire target frame images and adjacent frame images in the video sequence sample with an interval time not exceeding a preset time threshold.
[0126] The extraction unit 302 is used to extract the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image.
[0127] In some embodiments, the extraction unit 302 is configured to: perform full-convolution processing on the target frame image and the adjacent frame images to obtain corresponding first image features and second image features; perform convolution processing on the first image features through a convolution layer with a preset size to obtain first conditional parameter information corresponding to the instance information in the first image features; perform convolution processing on the second image features through a convolution layer with a preset size to obtain second conditional parameter information corresponding to the instance information in the second image features.
[0128] The input unit 303 is configured to input the first conditional parameter information and the second conditional parameter information into a preset model, and perform cross-mask recognition training on the target frame image and the adjacent frame images to obtain the trained preset model.
[0129] In some embodiments, the input unit 303 includes:
[0130] The first processing subunit is configured to perform mask convolution processing on the first image features to obtain first mask feature information;
[0131] The second processing subunit is configured to perform mask convolution processing on the second image features to obtain second mask feature information;
[0132] The input subunit is configured to input the first conditional parameter information and the second conditional parameter information into a preset model, and perform cross-mask recognition training on the first mask feature information and the second mask feature information to obtain the trained preset model.
[0133] In some embodiments, the input subunit is configured to: input the first conditional parameter information and the second conditional parameter information into a preset model; perform convolution processing on the first mask feature information and the second mask feature information respectively through the first conditional parameter information to obtain first mask information and first cross-mask information; perform convolution processing on the first mask feature information and the second mask feature information respectively through the second conditional parameter information to obtain second mask information and second cross-mask information; calculate a first difference value between the first mask information, the first cross-mask information, the second mask information and the second cross-mask information and the corresponding label information; iteratively adjust the network model parameters of the preset model according to the first difference value until the first difference value converges to obtain the trained preset model.
[0134] The recognition unit 304 is configured to perform mask recognition processing on the video sequence to be recognized based on the trained preset model.
[0135] In some embodiments, the determination unit 304, the recognition unit, is configured to: obtain the image to be recognized in each frame of the video sequence to be recognized; extract the third conditional parameter information corresponding to the instance information in the image to be recognized; input the third conditional parameter information into the trained preset model, and output the corresponding target mask information.
[0136] In some embodiments, the image processing device further includes:
[0137] A sample recognition unit, configured to recognize the sample mask information corresponding to the instance information in the sample image of each frame in the video sequence sample based on the trained preset model;
[0138] A classification training unit, configured to input the sample mask information into a preset classification model for classification training to obtain a trained preset classification model.
[0139] In some embodiments, the classification training unit is configured to: input the sample mask information into a preset classification model; convert the sample mask information into a vector representation space of a corresponding preset dimension; classify the instance information in the sample image of each frame according to the similarity between the vector representation spaces; and iteratively adjust the classification network parameters in the preset classification model according to the second difference value between the classification result and the label classification result until the second difference value converges, so as to obtain a trained preset classification model.
[0140] In some embodiments, the image processing device further includes: a classification unit, configured to classify the instance information in the video sequence to be recognized according to the trained preset classification model, so as to implement instance segmentation of the video sequence to be recognized.
[0141] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated herein.
[0142] As described above, in the embodiment of the present application, the acquisition unit 301 acquires the image to be processed, and acquires consecutive target frame images and adjacent frame images in the video sequence sample; the extraction unit 302 extracts the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image; the input unit 303 inputs the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtains the trained preset model; the recognition unit 304 performs mask recognition processing on the video sequence to be recognized based on the trained preset model. In this way, adjacent target frame images and adjacent frame images in the video sequence sample can be obtained, and the corresponding conditional parameter information of both can be extracted as convolution kernels respectively to perform cross-mask recognition learning on them, and a trained preset model with more accurate mask recognition can be obtained for recognition. Compared with the solution of performing image recognition and segmentation on a single-frame image, the embodiment of the present application can learn cross-frame image information, thereby greatly improving the accuracy of image processing.
[0143] The embodiment of the present application also provides a server, as Figure 6 shown, which shows the structural schematic diagram of the server involved in the embodiment of the present application. Specifically:
[0144] The server may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable computer storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 6 the server structure shown in
[0145] does not constitute a limitation on the server, and may include more or fewer components than shown in the figure, or combine certain components, or arrange different components. Among them:
[0146] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the server. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0147] The server further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0148] The server may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0149] Although not shown, the server may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the server will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:
[0150] Obtain consecutive target frame images and adjacent frame images in a video sequence sample; extract first conditional parameter information corresponding to instance information in the target frame image and second conditional parameter information corresponding to instance information in the adjacent frame image; input the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtain a trained preset model; perform mask recognition processing on the video sequence to be recognized based on the trained preset model.
[0151] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the detailed description of the image processing method above, which will not be elaborated here.
[0152] As can be seen from the above, the server in the embodiment of the present application can obtain the image to be processed, and obtain consecutive target frame images and adjacent frame images in the video sequence sample; extract the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image; input the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtain the trained preset model; perform mask recognition processing on the video sequence to be recognized based on the trained preset model. In this way, consecutive target frame images and adjacent frame images in the video sequence sample can be obtained, and the corresponding conditional parameter information of both can be extracted as convolution kernels respectively to perform cross-mask recognition learning on them, and a trained preset model with more accurate mask recognition is obtained for recognition. Compared with the solution of performing image recognition and segmentation on a single frame image, the embodiment of the present application can learn cross-frame image information, thereby greatly improving the accuracy of image processing.
[0153] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable computer storage medium and loaded and executed by a processor.
[0154] For this reason, the embodiment of the present application provides a computer storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiment of the present application. For example, the instructions can execute the following steps:
[0155] Obtain consecutive target frame images and adjacent frame images in the video sequence sample; extract the first conditional parameter information corresponding to the instance information in the target frame image and the second conditional parameter information corresponding to the instance information in the adjacent frame image; input the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-mask recognition training on the target frame image and the adjacent frame image, and obtain the trained preset model; perform mask recognition processing on the video sequence to be recognized based on the trained preset model.
[0156] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners provided in the above embodiments.
[0157] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, and details will not be repeated here.
[0158] Among them, the computer storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0159] Since the instructions stored in the computer storage medium can execute the steps in any one of the image processing methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the image processing methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0160] The above has introduced in detail an image processing method, apparatus, device and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain consecutive target frame images and adjacent frame images in a video sequence sample; Extract a first image feature and first conditional parameter information corresponding to instance information in the target frame image, and a second image feature and second conditional parameter information corresponding to instance information in the adjacent frame image, and perform masked convolution processing on the first image feature and the second image feature respectively to obtain first masked feature information and second masked feature information; Perform convolution processing on the first masked feature information and the second masked feature information respectively through the first conditional parameter information to obtain first masked information and first cross-masked information; Perform convolution processing on the first masked feature information and the second masked feature information respectively through the second conditional parameter information to obtain second masked information and second cross-masked information; Calculate a first difference value between the first masked information, the first cross-masked information, the second masked information, and the second cross-masked information and corresponding label information; Iteratively adjust the network model parameters of a preset model according to the first difference value until the first difference value converges to obtain a trained preset model; Perform masked recognition processing on an unknown video sequence based on the trained preset model.
2. The image processing method according to claim 1, wherein The extracting the first conditional parameter information corresponding to instance information in the target frame image and the second conditional parameter information corresponding to instance information in the adjacent frame image includes: Perform fully convolutional processing on the target frame image and the adjacent frame image to obtain corresponding first image features and second image features; Perform convolution processing on the first image feature through a convolutional layer with a preset size to obtain first conditional parameter information corresponding to instance information in the first image feature; Perform convolution processing on the second image feature through a convolutional layer with a preset size to obtain second conditional parameter information corresponding to instance information in the second image feature.
3. The image processing method according to claim 2, wherein Inputting the first conditional parameter information and the second conditional parameter information into a preset model to perform cross-masked recognition training on the target frame image and the adjacent frame image to obtain a trained preset model includes: Perform masked convolution processing on the first image feature to obtain first masked feature information; Perform masked convolution processing on the second image feature to obtain second masked feature information; Input the first conditional parameter information and the second conditional parameter information into the preset model to perform cross-masked recognition training on the first masked feature information and the second masked feature information to obtain a trained preset model.
4. The image processing method according to any one of claims 1 to 3, characterized in that The obtaining consecutive target frame images and adjacent frame images in a video sequence sample includes: Obtain target frame images and adjacent frame images in a video sequence sample with an interval time not exceeding a preset time threshold.
5. The image processing method according to any one of claims 1 to 3, characterized in that, The performing masked recognition processing on an unknown video sequence based on the trained preset model includes: Obtain an unknown image for each frame in the unknown video sequence; Extract third conditional parameter information corresponding to instance information in the unknown image; Input the third conditional parameter information into the trained preset model to output corresponding target masked information.
6. The image processing method according to any one of claims 1 to 3, characterized in that, The image processing method further includes: Based on the trained preset model, identify the sample mask information corresponding to the instance information in each frame of the sample image in the video sequence sample; Input the sample mask information into a preset classification model for classification training to obtain the trained preset classification model.
7. The image processing method according to claim 6, wherein The step of inputting the sample mask information into a preset classification model for classification training to obtain the trained preset classification model includes: Input the sample mask information into a preset classification model; Convert the sample mask information into a vector representation space of a corresponding preset dimension; Classify the instance information in each frame of the sample image according to the similarity between the vector representation spaces; Iteratively adjust the classification network parameters in the preset classification model according to the second difference value between the classification result and the labeled classification result until the second difference value converges to obtain the trained preset classification model.
8. The image processing method according to claim 7, wherein The image processing method further includes: Classify the instance information in the video sequence to be recognized according to the trained preset classification model to implement instance segmentation of the video sequence to be recognized.
9. An image processing apparatus, characterized in that, It includes: An acquisition unit for acquiring consecutive target frame images and adjacent frame images in the video sequence sample; An extraction unit for extracting the first image feature and the first conditional parameter information corresponding to the instance information in the target frame image and the second image feature and the second conditional parameter information corresponding to the instance information in the adjacent frame image, and respectively performing mask convolution processing on the first image feature and the second image feature to obtain first mask feature information and second mask feature information; An input unit for respectively performing convolution processing on the first mask feature information and the second mask feature information through the first conditional parameter information to obtain first mask information and first cross-mask information; Respectively perform convolution processing on the first mask feature information and the second mask feature information through the second conditional parameter information to obtain second mask information and second cross-mask information; Calculate the first difference value between the first mask information, the first cross-mask information, the second mask information and the second cross-mask information and the corresponding label information; iteratively adjust the network model parameters of the preset model according to the first difference value until the first difference value converges to obtain the trained preset model; An identification unit for performing mask identification processing on the video sequence to be recognized based on the trained preset model.
10. The processing device according to claim 9, characterized in that, The extraction unit is used for: Performing full convolution processing on the target frame image and the adjacent frame image to obtain corresponding first image features and second image features; Performing convolution processing on the first image feature through a convolution layer of a preset size to obtain the first conditional parameter information corresponding to the instance information in the first image feature; Performing convolution processing on the second image feature through a convolution layer of a preset size to obtain the second conditional parameter information corresponding to the instance information in the second image feature.
11. The processing device according to claim 10, characterized in that, The input unit includes: A first processing subunit for performing mask convolution processing on the first image feature to obtain first mask feature information; A second processing subunit, configured to perform masked convolution processing on the second image feature to obtain second masked feature information; An input subunit, configured to input the first conditional parameter information and the second conditional parameter information into a preset model, and perform cross-masked recognition training on the first masked feature information and the second masked feature information to obtain a trained preset model.
12. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the image processing method according to any one of claims 1 to 8 are implemented.
13. A computer storage medium, characterized in that, The computer storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the image processing method according to any one of claims 1 to 8.