Training method of image processing model, and video image processing method and system

Through training the image processing model, using twin networks to process deformed images, the accuracy and stability problems of human skin area detection and segmentation in live videos are solved, and efficient, real-time and accurate target area segmentation is achieved.

CN120219912AActive Publication Date: 2025-06-27GUANGZHOU HUYA TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510236317.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, poor compatibility, poor real-time and low stability when detecting and segmenting human skin areas in live video broadcasts, especially when dealing with light changes, occlusion and complex backgrounds.

Method used

A training method of image processing model is adopted. By obtaining sample images, cropping target objects, triangulation deformation, marking segmentation objects, and inputting the marked samples and deformation samples into the twin network for training, improving the robustness of the twin network and the ability to process deformed images.

Benefits of technology

It realizes precise segmentation and processing of human skin areas in live video, improves segmentation accuracy and stability, enhances the processing ability of deformed images, and ensures real-time and efficient in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219912A_ABST
    Figure CN120219912A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to an image processing model training method, a video image processing method and a video image processing system. The method comprises the following steps: extracting a video frame image from a video as a to-be-processed image; cutting the to-be-processed image according to the region of interest to obtain a cut image; inputting the cut image into a twin network for processing to obtain a segmented region image; the twin network training method comprises the following steps: acquiring a plurality of sample images; performing target object cutting on the sample image to obtain a sample cutting image; performing triangulation deformation on the sample cutting image to obtain a corresponding deformed sample image; marking segmentation objects for the sample image and the corresponding deformed sample image; and respectively inputting the marked sample cutting image and the corresponding deformed sample image into a twin network, and training the twin network to obtain a trained twin network. According to the method, accurate segmentation and processing of the segmentation object in the image can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more particularly, to a method for training an image processing model, a method and a system for video image processing. Background Art

[0002] In the field of video live broadcast, accurately detecting and segmenting the target area is a key technology, especially the human skin area; the detected and segmented human skin area can be used for portrait beauty, including multiple functions such as removing spots and acne, beauty retouching, and skin tone adjustment. However, the current existing technologies still have many deficiencies in terms of accuracy, compatibility, real-time performance, and stability: the algorithms used in the existing technologies perform poorly when dealing with light changes, occlusions, and complex backgrounds in pictures, and cannot accurately identify and segment the target area; the target segmentation technology of the current existing technologies is difficult to meet the real-time requirements in practical applications; when segmenting the skin area with the existing technologies, it is still a great challenge to process different ethnic groups, skin colors, and environmental illuminations, and there are often large errors in identifying and segmenting the skin area. Therefore, it is necessary to make necessary improvements to the technology for detecting and segmenting the target area. Summary of the Invention

[0003] The present invention aims to overcome at least one defect (deficiency) of the above-mentioned existing technologies, and provides a method for training an image processing model, a method and a system for video image processing, which are used to realize the accurate segmentation and processing of the segmentation object in the image.

[0004] According to the first aspect of the present application, there is provided a method for training an image processing model, the method comprising:

[0005] Obtaining a plurality of sample images;

[0006] Performing target object cropping on the sample images to obtain sample cropped images;

[0007] Performing triangulation deformation on the sample cropped images to obtain corresponding deformed sample images;

[0008] Marking the segmentation object for the sample images and the corresponding deformed sample images;

[0009] Inputting the marked sample cropped images and the corresponding deformed sample images into a siamese network respectively, and training the siamese network to obtain a trained siamese network.

[0010] It can be understood that training the siamese network with the sample cropped images and the corresponding deformed sample images can be compatible with the images deformed due to uncontrollable reasons, so as to improve the robustness of the siamese network; at the same time, it can also solve the problem that the target area cannot be accurately segmented for subsequent videos or pictures due to deformation, and improve the reliability of the segmentation of videos or pictures.

[0011] Optionally, inputting the marked sample cropped image and the corresponding deformed sample image into the Siamese network respectively, training the Siamese network, and obtaining a trained Siamese network, including:

[0012] Inputting the marked sample cropped image into the Siamese network to obtain a first generated image, and inputting the deformed sample image corresponding to the sample cropped image into the Siamese network to obtain a corresponding second generated image;

[0013] Obtaining the loss value of the Siamese network according to the first generated image and the corresponding second generated image;

[0014] Optimizing the parameters of the Siamese network according to the loss value to obtain a trained Siamese network.

[0015] It can be understood that triangulating and deforming the sample image to obtain the corresponding deformed sample image, and then generating the second generated image from it actually simulates the deformation and distortion of the region where the segmented object is located in the video due to network problems or target object movement problems. And training the Siamese network with the deformed sample image can enhance the processing ability of the Siamese network for deformed images, so that when the Siamese network processes deformed images, it can also accurately segment the region where the segmented object is located, enhancing the robustness of the Siamese network.

[0016] Optionally, the obtaining the loss value of the Siamese network according to the first generated image and the corresponding second generated image includes:

[0017] Performing an inverse deformation operation on the second generated image to obtain an inversely deformed second generated image;

[0018] Performing a global pixel comparison between the first generated image and the inversely deformed second generated image to obtain the loss value of the Siamese network.

[0019] It can be understood that the loss value of the Siamese network is obtained according to the similarity between the first generated image and the inversely deformed second generated image, and the second generated image simulates the situation of using the Siamese network after the sample cropped image is deformed. Through this loss value, the Siamese network can segment the region where the segmented object in the deformed sample image is located, tending to the region where the segmented object in the non-deformed image is located, so as to achieve the effect of enhancing the ability of the Siamese network to process deformed images.

[0020] According to the second aspect of the present application, a video image processing method is provided, and the method includes:

[0021] Extracting video frame images from the video as images to be processed;

[0022] Crop the to-be-processed image according to the region of interest to obtain a cropped image;

[0023] Input the cropped image into a siamese network for processing to obtain a segmented region image; wherein the siamese network is a siamese network trained by using the training method of an image processing model described in the first aspect.

[0024] It can be understood that by using the trained siamese network to process video frame images, it is possible to stably and accurately segment and detect the region where the segmentation object is located even in the case of video frame images with complex backgrounds, improve the segmentation accuracy of the region where the segmentation object is located, and at the same time achieve efficient real-time operation of video processing. At the same time, it can be compatible with the processing of deformed images caused by network problems or human body movements in the video, enhance the versatility of the method, and ensure stable operation in different video situations.

[0025] Optionally, the region of interest is stored in a preset database;

[0026] Extract the specific position information of the target object in the to-be-processed image according to the segmented region image;

[0027] Generate a region of interest for the current video frame image according to the specific position information;

[0028] After replacing the region of interest of the current video frame image with the region of interest stored in the database, store it as the region of interest corresponding to the next video frame image.

[0029] It can be understood that the segmented region image precisely reflects the specific position of the segmentation object in the to-be-processed image, and can generate a region of interest for the current video frame image according to the specific position information of the segmentation object, providing guidance for the cropping work of the next video frame image, thereby effectively narrowing the range of target object segmentation and improving the processing efficiency of the siamese network; by dynamically updating the region of interest in the database, it is possible to more precisely intercept a small-range cropped image of the video frame image, thereby significantly improving the true resolution of the cropped image input to the siamese network and reducing the consumption of computing resources; at the same time, it also ensures the continuity of the target in the complex environment of the video, reduces the interference of background objects on the segmentation result of the segmentation object, and improves the robustness of segmentation and detection.

[0030] Optionally, the method further includes:

[0031] According to the segmented region image corresponding to the previous video frame image, obtain the motion vector of each pixel in the segmented region image corresponding to the current video frame image through the fast optical flow algorithm;

[0032] Adjust the noise or incoherent regions in the image of the current segmentation region corresponding to the current video frame image according to the motion vector to obtain the processed current segmentation region image.

[0033] It can be understood that by performing a smoothing operation on the segmentation region image corresponding to the current video frame image through the fast optical flow algorithm, the change of the segmentation region in the continuous video frame images in the video can be effectively detected, the short-term motion law can be introduced into the static segmentation of the video frame image, the jitter and mutation of the segmentation region of the segmentation object can be significantly reduced, the noise or incoherent regions caused by segmentation errors can be reduced, and the stability and continuity of the segmentation result can be improved.

[0034] Optionally, the method further includes:

[0035] Expand the boundary of the segmentation region image by applying a morphological dilation operation;

[0036] Perform feathering processing on the boundary to obtain a segmentation region image with a faded boundary region;

[0037] Perform Gaussian blur filtering on the segmentation region image with the faded boundary region to obtain the final segmentation region image.

[0038] It can be understood that through morphological dilation, feathering processing and Gaussian blur filtering, by expanding the boundary of the target segmentation region and performing processing, the edge of the region where the segmentation object is located in the segmentation region image can be made to transition more smoothly and naturally. When applied to the segmentation of human skin, it helps to handle situations such as fine hairs and blurred edges, and achieves a more natural skin segmentation effect.

[0039] According to the third aspect of the present application, a training system for an image processing model is provided. The system includes:

[0040] A sample acquisition module for acquiring a number of sample images;

[0041] A sample cropping module for cropping the target object from the sample image to obtain a sample cropped image;

[0042] A deformation processing module for triangulating and deforming the sample cropped image to obtain a corresponding deformed sample image;

[0043] A marking module for marking the segmentation object of the sample image and the corresponding deformed sample image;

[0044] A training module for respectively inputting the marked sample cropped image and the corresponding deformed sample image into a siamese network, training the siamese network, and obtaining a trained siamese network.

[0045] According to a fourth aspect of the present application, a video image processing system is provided, and the system includes:

[0046] An acquisition module, configured to extract video frame images from a video as images to be processed;

[0047] A cropping module, configured to crop the image to be processed according to a region of interest to obtain a cropped image;

[0048] A generation module, configured to input the cropped image into a siamese network for processing to obtain a segmented region image; wherein the siamese network is a siamese network trained by using the training method of an image processing model described in the first aspect.

[0049] According to a fifth aspect of the present application, an electronic device is provided, including:

[0050] A memory, configured to store one or more computer programs;

[0051] A processor, when the one or more computer programs are executed by the processor, implementing the training method of the image processing model described in the first aspect or a video image processing method described in the second aspect above.

[0052] According to a sixth aspect of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the training method of the image processing model described in the first aspect or a video image processing method described in the second aspect above when executed by a processor.

[0053] Based on any one of the above aspects, a training method of an image processing model, a video image processing method, a system, an electronic device, and a storage medium provided in an embodiment of the present application, by extracting video frame images from a video as images to be processed; cropping the image to be processed according to a region of interest to obtain a cropped image; inputting the cropped image into a siamese network for processing to obtain a segmented region image; wherein the training of the siamese network includes: acquiring a plurality of sample images; performing target object cropping on the sample images to obtain sample cropped images; performing triangulation deformation on the sample cropped images to obtain corresponding deformed sample images; performing segmentation object marking on the sample images and the corresponding deformed sample images; inputting the marked sample cropped images and the corresponding deformed sample images into the siamese network respectively, and training the siamese network to obtain a trained siamese network. The method can bring the following benefits:

[0054] 1) Training the siamese network according to the sample cropped images and the corresponding deformed sample images can enhance the ability of the siamese network to segment segmentation objects from deformed images, and improve the reliability and robustness of the image processing model based on the siamese network;

[0055] 2) In the field of video processing, especially in the scenario of video live streaming, an efficient, stable, accurate and real-time target area segmentation method is provided, which has good general performance;

[0056] 3) It can well solve the problems of poor accuracy and poor inter-frame stability for videos, especially video live streaming. By constructing the structure of a Siamese network, the training loss optimized for stability is increased, and the problem of the decrease in the accuracy of the segmented target area due to image distortion during video live streaming is optimized, avoiding the possibility of a large gap in the target segmentation area between two frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0058] Figure 1 Schematic diagram of the application scenario of a training method for an image processing model and a video image processing method provided in this embodiment.

[0059] Figure 2 Flowchart of a training method for an image processing model provided in this embodiment.

[0060] Figure 3 Flowchart of a Siamese network training method provided in this embodiment.

[0061] Figure 4 Flowchart of a video image processing method provided in this embodiment.

[0062] Figure 5 Flowchart of a method for obtaining an area of interest provided in this embodiment.

[0063] Figure 6 Flowchart of a method for smoothing the segmented area image provided in this embodiment.

[0064] Figure 7 Flowchart of a method for feathering the segmented area image provided in this embodiment.

[0065] Figure 8 Schematic diagram of the modules of a training method for an image processing model provided in this embodiment.

[0066] Figure 9 Schematic diagram of the modules of a video image processing system provided in this embodiment.

[0067] Figure 10 Schematic structural diagram of the electronic device provided in this embodiment. Detailed implementation manners

[0068] The accompanying drawings of this application are only for illustrative purposes and should not be construed as a limitation to this application. To better illustrate the following embodiments, some components in the drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0069] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0070] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0071] Video live broadcast refers to a way of real-time live broadcast using the Internet and streaming media technology. At present, video live broadcast has been applied to multiple fields and develops rapidly. In order to present a better live broadcast effect, live broadcasters often use the beautification effects provided by live broadcast software, such as beauty functions like freckle and acne removal, beauty retouching, and skin tone adjustment. For a live broadcast software to present better beauty function effects, it needs to return to a fundamental problem: whether it can accurately identify the skin area of the human body during the live broadcast, especially the face of the live broadcaster. Nowadays, although many existing technologies have been disclosed that can segment the skin area of the human body, however, in the face of problems such as complex live broadcast backgrounds, diverse ethnic skin colors, dynamic shaking of the human body movements during the live broadcast, and picture distortion caused by network problems and lags, the skin areas segmented in video live broadcast often have large errors and noise, and the beauty effect also decreases or the beauty object is inaccurate. Therefore, it is necessary to make necessary improvements to the segmentation and processing of the target area, especially the skin area.

[0072] This embodiment provides a technical solution that can solve the above problems. Next, in conjunction with the accompanying drawings, the specific implementation manners of the present application will be described in detail.

[0073] Exemplarily, it is a schematic diagram of an application scenario of a training method for an image processing model and a video image processing method provided in an embodiment of the present application. As Figure 1 shown, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a function of training a twin network; the terminal device 200 has a video playback function and can also have an image acquisition function.

[0074] It can be understood that the server 100 can be an independent electronic device or a cluster composed of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.

[0075] In an implementable manner, the server 100 and the terminal 200 can respectively execute the training method for the image processing model or the video image processing method provided in the embodiment of the present application. Alternatively, preferably, part of the training method for the image processing model or the video image processing method provided in the embodiment of the present application is executed in the server 100, and part is executed in the terminal 200.

[0076] Exemplarily, as Figure 2 shown, this embodiment provides a training method for an image processing model, which may include the following steps:

[0077] S110. Obtain a plurality of sample images;

[0078] In this embodiment, a pre-collected historical video set or picture set can be used as the sample images. The picture set can include a large number of representative situations of picture deformation or distortion. Preferably, some pictures are individually discontinuous, and data enhancement is required for the pictures to highlight features, so that the twin network can more easily learn the processing methods of a large number of deformed and distorted images, thereby improving the robustness of the twin network.

[0079] S120. Crop the target object from the sample images to obtain sample cropped images;

[0080] In this embodiment, cropping the target object from the sample image can narrow the calculation range of the siamese network for the sample image, thereby improving the segmentation efficiency of the siamese network and saving computing resources. The recognition of the target object is completed before cropping, which can be achieved through manual recognition or machine recognition in the prior art and will not be elaborated here. The target object can be a specific type of object, such as a person, an animated character, a digital human, etc.

[0081] S130. Perform triangulation deformation on the sample cropped image to obtain a corresponding deformed sample image.

[0082] In this embodiment, some of the collected sample images may be normal and need to be deformed to enhance their deformation and distortion characteristics. Triangulation deformation constructs a triangular mesh based on the pixel point set of the image, divides the image into multiple triangular regions, and is commonly used as a deformation operation on images in the field of image processing. Using the method of triangulation deformation can process the picture using simple mathematical and geometric principles, enabling the siamese network to more easily recognize the features therein and improving the processing efficiency of the siamese network.

[0083] S140. Mark the segmentation objects for the sample image and the corresponding deformed sample image.

[0084] In this embodiment, when training the siamese network, it is necessary to pre-mark the segmentation objects in the sample image to form training samples with labels, so that the siamese network can learn the method of segmenting the segmentation objects from the training samples with labels, forming a supervised machine learning process. The segmentation object is a part of the target object. Training the siamese network aims to perform segmentation recognition on the image through the siamese network, so as to segment the area where the segmentation object is located in the image, and then subsequent processing can be performed on the area where the segmentation object is located in the image. Exemplarily, when the target object is a person, the segmentation object can be the skin, eyes, nose, etc. of the person.

[0085] S150. Input the marked sample cropped image and the corresponding deformed sample image into the siamese network respectively, and train the siamese network to obtain a trained siamese network.

[0086] In this embodiment, it is a necessary function of the method to satisfy the interception of the area where the segmentation object of the normal image is located. However, in the application scenario of live video, due to the real-time movement of people in the video or the picture distortion caused by network problems such as lag, the area where the segmented object is located often has other background pictures or incomplete areas, resulting in excessive or insufficient rendering of the area to be processed by the subsequent beauty function, and the live effect is unnatural. This application not only uses the normal cropped image as the training data of the siamese network, but also artificially performs triangulation deformation on the normal cropped image to simulate the distortion of the area where the segmentation object is located in the real situation, so as to train the siamese network's cropping ability for distorted pictures at the same time. Because there are strong requirements for the accuracy and stability of the target area segmentation in video live broadcast, this application uses a siamese network and trains it to obtain a trained target area segmentation model. Preferably, DeepLabV3+ is selected as the basic deep learning framework, and MobileNetV3 is selected as the underlying model, and two branches with shared parameters are built using the siamese network.

[0087] In this embodiment, the labeled sample cropped image and the corresponding deformed sample image actually simulate the situation after the sample cropped image is deformed and distorted. Inputting the sample cropped image and the corresponding deformed sample image into the siamese network respectively can enable the siamese network to learn the segmentation method of the sample cropped image in the normal situation for the area where the segmentation object is located, and the segmentation method of the corresponding deformed sample image in the deformed situation for the area where the segmentation object is located, enhancing the reliability and generality of the siamese network.

[0088] Specifically, as Figure 3 shown, the steps of inputting the labeled sample cropped image and the corresponding deformed sample image into the siamese network respectively, training the siamese network, and obtaining a trained siamese network include the following steps:

[0089] S151. Input the labeled sample cropped image into the siamese network to obtain a first generated image, and input the deformed sample image corresponding to the sample cropped image into the siamese network to obtain a corresponding second generated image;

[0090] In this embodiment, a normal sample cropped image is input into the Siamese network, which reflects the area where the segmentation object is located under normal circumstances. It is understandable that the first generated image output by this input branch ultimately needs to be trained to generate a result that conforms to the area where the segmentation object is located; a corresponding deformed sample image is input into the Siamese network, which reflects the area where the segmentation object is located under the condition of a distorted picture. It is understandable that the second generated image output by this input branch ultimately needs to be trained to be close to the first generated image, so as to illustrate that the Siamese network can also accurately segment the area where the segmentation object is located in the distorted picture.

[0091] In this embodiment, the sample cropped image and the corresponding deformed sample image are respectively input into the Siamese network, which can enable the Siamese network to extract the features in the two input branches respectively, and learn the information and methods in the two input branches separately, so as to avoid the situation of information confusion and cross-disorder processing between the two input branches.

[0092] S152. Obtain the loss value of the Siamese network according to the first generated image and the corresponding second generated image;

[0093] In this embodiment, training the Siamese network to process distorted pictures is a continuous optimization process. By comparing the similarity between the first generated image and the corresponding second generated image, the difference between the two is obtained, and the loss value of the Siamese network is generated according to the difference. According to the loss value, the parameters and structure of the Siamese network are fine-tuned and optimized, so that the two images obtained by training the two branches next time are continuously approaching similarity, thereby improving the ability of the Siamese network to process distorted images.

[0094] Specifically, the obtaining of the loss value of the Siamese network according to the first generated image and the corresponding second generated image includes:

[0095] Perform an inverse deformation operation on the second generated image to obtain an inversely deformed second generated image;

[0096] Perform a global pixel comparison between the first generated image and the inversely deformed second generated image to obtain the loss value of the Siamese network.

[0097] Preferably, the loss value Loss is specifically obtained by the following formula:

[0098] Loss = (F L (Net ori (img), F rewarp (Net warp (img warp )))

[0099] where img represents the sample cropped image, Netori (img) represents the first generated image obtained by inputting the cropped sample image into the siamese network; img warp represents the deformed sample image corresponding to the cropped sample image, Net warp (img warp ) represents the second generated image obtained by inputting the deformed sample image corresponding to the cropped sample image into the siamese network; F rewarp F(·) is the inverse deformation operation function for the image, F L F(·) is the global pixel contrast operation function for the first generated image and the second generated image.

[0100] In this embodiment, the deformed sample image obtained by autonomously performing triangulation deformation on the cropped sample image is further processed by the siamese network to generate a processed image. An inverse processing of the deformation is required to restore the generated processed image to its normal state for better comparison with the first generated image. Through the contrast operation, each pixel of the first generated image and the second generated image is compared, which can improve the accuracy of obtaining the differences.

[0101] S153. Optimize the parameters of the siamese network according to the loss value to obtain a trained siamese network.

[0102] In this embodiment, by setting a threshold for stopping optimization, when the loss value reaches the threshold for stopping optimization, the optimization is stopped. It can be understood that the obtained sample images are usually divided into a training set and a test set. The training of the siamese network usually includes a training process and a test process. After optimizing the siamese network using the loss value of the training set and completing the training operation, the test set needs to be input into the siamese network to test the siamese network. After completing the test, a trained siamese network is obtained. It can be understood that the trained siamese network can process normal cropped images and distorted cropped images in actual situations and can be directly used in the subsequent segmentation steps for segmenting objects.

[0103] As Figure 4 shown, an embodiment of the present application further provides a video image processing method, and the method includes:

[0104] S210. Extract video frame images from the video as images to be processed;

[0105] Preferably, in the live video, the video frame image to be processed is extracted as the image to be processed, and each video frame image can be extracted, so that the segmentation of the segmentation object can be performed on each video frame image. For the live video, it is usually necessary to perform beauty treatment on the people in it. The area to be beautified can be used as the segmentation object. For example, if the skin needs to be beautified, the video image processing method of the present application is used to segment the area to be beautified, so that the subsequent beauty treatment can be more natural.

[0106] S220. Crop the image to be processed according to the region of interest to obtain a cropped image;

[0107] In this embodiment, considering that the background image in the live video scenario is complex and the person as the target object may be relatively small in the picture proportion, in order to better exclude the interference of the background in the video image and other non-target objects, small portraits, etc. on the effect, the present application adopts a pre-processing method of image cropping. Obtain the video frame image to be processed in the live video for use as the material for the recognition and segmentation of the segmentation object; it can be understood that the region of interest is generated according to the specific target position information obtained from the segmentation region image of the previous video frame image; it can be understood that cropping the region where the target object of the current video frame image is located according to the region of interest can reduce the access to the recognition and calculation of the segmentation object in the subsequent process, reduce the interference of the large background area, and thus improve the accuracy of the recognition of the region where the segmentation object is located; at the same time, the region of interest recorded from the previous video frame image can ensure the continuity and stability of the region where the target object and / or segmentation object is located in the video to a certain extent. Because generally, the movement range of the live broadcaster in two adjacent frames is not large, and the position of the region where the target object and / or segmentation object is located in two adjacent frames will not move too obviously, so the region where the target object and / or segmentation object is located between two frames has continuity. When the segmentation image result is applied to the beauty treatment, it can ensure that the recognized region where the target object and / or segmentation object is located has continuity, so that the subsequent beauty function processing is more natural;

[0108] It can be understood that at the start time of processing, the region of interest in the database is empty, and the interception of the video frame image cannot be guided by the region of interest. Therefore, the entire video frame image size is cropped, that is, the cropping box coordinates are [0, 0, height, width]. More computing resources need to be obtained to identify the target region to obtain the accurate target segmentation region; but after obtaining the segmentation region image of the first video frame image, the corresponding region of interest can be generated according to the segmentation region image and stored by replacing the empty data in the database to guide the cropping work of the next frame.

[0109] In this embodiment, the coordinates of the cropping frame are generated according to the region of interest, and the corresponding cropping frame coordinates are obtained from the video frame image in advance, which facilitates the subsequent cropping work.

[0110] In this embodiment, the guiding region img mapped to the current video frame image img is obtained according to the region of interest a , and the specific formula for obtaining it is as follows:

[0111] img a = img[x left :x right ,y top :y bottom

[0112] where x left , x right , y top and y bottom are respectively the leftmost coordinate point, the rightmost coordinate point, the uppermost coordinate point and the lowermost coordinate point of the region of interest mapped to the current video frame image img;

[0113] The image to be processed is cropped according to the guiding region img a to obtain a cropped image;

[0114] It can be understood that the region of interest can guide the cropping work. It can be understood that the target object is a part of the region of interest. The region of interest can locate the specific approximate position of the segmentation object in the current video frame image. However, the segmentation object in the next video frame image will have slight differences compared with the current video frame image. Therefore, the cropping work of the next video frame image can only use the region of interest as a reference, and the cropped image obtained by guiding the cropping according to the region of interest is an image containing the entire target object.

[0115] Specifically, as Figure 5 shown, the region of interest is stored in a preset database; the steps for obtaining the region of interest include the following:

[0116] S221. Extract the specific position information of the target object in the image to be processed according to the segmentation region image;

[0117] In this embodiment, the region of interest ROI (Region of Interest) is a certain target region or target object in the image that we are particularly concerned about. These regions usually contain key position information that needs to be recognized, analyzed or processed.

[0118] S222. Generate the region of interest of the current video frame image according to the specific position information;

[0119] ​In this embodiment, an interested region is generated based on the segmentation region image results obtained from each video frame image, and the interested region in the database is updated, which can provide an accurate ROI (Region of Interest) for subsequent cropping of video frame images, reduce the computational amount of subsequent video frame images, and improve the stability and accuracy of target tracking.

[0120] S223. After replacing the interested region of the current video frame image with the interested region stored in the database, store it for use as the interested region corresponding to the next video frame image.

[0121] In this embodiment, the database always stores the interested region obtained from the previous video frame image. Before the start of the processing work, the data in the database is empty; when the current interested region is obtained, it is necessary to replace the interested region obtained from the previous video frame image in the database, overwrite and write the current interested region, making the storage form and content simpler, without the need to store the correspondence between the interested region and the video frame image, saving memory at the same time, and reducing the complexity of the corresponding processing work. For different videos, specifically in implementation, before processing different videos, the interested region in the database can be cleared in advance to avoid interference of the target object in the previous video on the target object in the next video.

[0122] S230. Input the cropped image into the Siamese network for processing to obtain a segmentation region image; where the Siamese network is the Siamese network trained by the above training method.

[0123] In this embodiment, for the trained Siamese network, only one input branch is needed to process the cropped image to obtain the segmentation region image.

[0124] In this embodiment, there are usually certain errors in the segmentation region image obtained through the Siamese network, especially at the boundary part of the region, there are a little noise points and serrated segmentation traces, making the entire live broadcast screen or beauty effect unnatural and affecting the user viewing experience. Therefore, it is necessary to post-process the obtained segmentation region image.

[0125] Specifically, as Figure 6 shown, the post-processing of the segmentation region image may include the following steps:

[0126] S241. According to the segmentation region image corresponding to the previous video frame image, obtain the motion vector of each pixel in the segmentation region image corresponding to the current video frame image through the fast optical flow algorithm;

[0127] In this embodiment, the fast optical flow algorithm is based on the basic assumptions of the optical flow method, including the constant brightness assumption, the time continuity assumption, and the spatial consistency assumption. It estimates the motion of the object by analyzing the changes in pixel intensity between consecutive video frame images. The core is to use the information between adjacent frame images to calculate the motion vector of each pixel point that changes with time. In this embodiment, it is used to calculate the motion vector of each pixel point in adjacent frame images. In this application, the fast optical flow algorithm is applied to the post-processing of the segmentation area results, which can effectively track the changes in the segmentation area images in the continuous video frame images, introduce the short-term motion law into the static segmentation, and significantly reduce the jitter and mutation of the segmentation area image.

[0128] S242: adjusting the noise points or incoherent areas in the segmented region image corresponding to the current video frame image according to the motion vector to obtain a processed current segmented region image.

[0129] In this embodiment, the motion vector adjusts the current segmented area image, which can reduce noise or incoherent areas caused by segmentation processing errors and improve the stability and continuity of the segmentation results. Preferably, the application uses a fast optical flow algorithm to calculate the optical flow field between the current video frame image and the previous video frame image for smoothing the current segmented area image.

[0130] Specifically, Figure 7 As shown, the post-processing of the segmented region image may further include the following steps:

[0131] S251, expanding the boundary of the segmented region image by applying a morphological dilation operation;

[0132] In this embodiment, by applying a morphological dilation operation to expand the boundary of the segmented region image, the segmented region image can be slightly enlarged, and the noise points and uneven areas in the boundary can be magnified.

[0133] S252, feathering the boundary to obtain a segmented area image with a faded boundary area;

[0134] In this embodiment, the feathering expansion technology is used to expand the boundary of the segmented region image, so that the boundary transition of the segmented region image is smoother and more natural, especially for the segmentation of the skin segmented region image, it is helpful to handle small hairs, blurred edges, etc., and achieve a more natural segmentation effect. In this embodiment, the expanded boundary area is feathered to make the boundary transition area gradually fade, thereby reducing the hardness of the boundary.

[0135] S253 , performing Gaussian blur filtering on the segmented region image with the faded boundary region to obtain a final segmented region image.

[0136] In this embodiment, Gaussian blur is an image processing technique that smooths an image by applying a normal distribution as a weight function, reducing the noise and details in the image and highlighting the main features. In this embodiment, the boundary transition region is further smoothed by the filtering method of Gaussian blur to ensure the naturalness and consistency of the boundary.

[0137] As Figure 8 shown, the embodiment of the present application also provides a training system for an image processing model. Optionally, the training system may include:

[0138] A sample acquisition module 311, a sample cropping module 312, a deformation processing module 313, a marking module 314, and a training module 315, where:

[0139] The sample acquisition module 311 is used to acquire a number of sample images;

[0140] In this embodiment, the sample acquisition module 311 can be used to execute Figure 2 the steps S110 shown, and the specific description of the sample acquisition module 311 can refer to the description of the steps S110.

[0141] The sample cropping module 312 is used to crop the target object from the sample image to obtain a sample cropped image;

[0142] In this embodiment, the sample cropping module 312 can be used to execute Figure 2 the steps S120 shown, and the specific description of the sample cropping module 312 can refer to the description of the steps S120.

[0143] The deformation processing module 313 is used to perform a triangulation deformation on the sample cropped image to obtain a corresponding deformed sample image;

[0144] In this embodiment, the deformation processing module 313 can be used to execute Figure 2 the steps S130 shown, and the specific description of the deformation processing module 313 can refer to the description of the steps S130.

[0145] The marking module 314 is used to mark the segmentation objects for the sample image and the corresponding deformed sample image;

[0146] In this embodiment, the marking module 314 can be used to execute Figure 2 the steps S140 shown, and the specific description of the marking module 314 can refer to the description of the steps S140.

[0147] A training module 315, configured to respectively input the marked sample cropped images and the corresponding deformed sample images into a siamese network, train the siamese network, and obtain a trained siamese network.

[0148] In this embodiment, the training module 315 can be used to execute Figure 2 the step S150 shown in the figure. For the specific description of the training module 315, reference can be made to the description of the step S150.

[0149] As Figure 9 shown in the figure, an embodiment of the present application further provides a video image processing system. Optionally, the system may include:

[0150] An acquisition module 411, a cropping module 412, and a generation module 413, where:

[0151] The acquisition module 411 is configured to extract video frame images from a video as images to be processed;

[0152] In this embodiment, the acquisition module 411 can be used to execute Figure 4 the step S210 shown in the figure. For the specific description of the acquisition module 411, reference can be made to the description of the step S210.

[0153] The cropping module 412 is configured to crop the image to be processed according to a region of interest to obtain a cropped image;

[0154] In this embodiment, the cropping module 412 can be used to execute Figure 4 the step S220 shown in the figure. For the specific description of the cropping module 412, reference can be made to the description of the step S220.

[0155] The generation module 413 is configured to input the cropped image into a siamese network for processing to obtain a segmented region image; where the siamese network is a siamese network trained by using a training method for an image processing model;

[0156] In this embodiment, the generation module 413 can be used to execute Figure 4 the step S230 shown in the figure. For the specific description of the generation module 413, reference can be made to the description of the step S230.

[0157] An embodiment of the present application further provides an electronic device, the structure of which is as Figure 10 shown in the figure. The electronic device includes a memory 511, a processor 512, a communication module 513, an input / output interface 514, etc. Optionally, the memory 511, the processor 512, the communication module 513, and the input / output interface 514 can be connected and communicate through a bus 515.

[0158] The memory 511 is used to store one or more computer programs and transfer the code of the computer program to the processor 512; when the one or more computer programs are executed by the processor 512, a training method for an image processing model or a video image processing method in an embodiment of the present application is implemented.

[0159] Optionally, the electronic device can be connected to a network via a communication module 513 to communicate with other devices, such as a terminal or a server, through the network to achieve data interaction. The electronic device can be various forms of digital computers, such as desktop computers, servers, workstations, mainframe computers or other types of computers. The electronic device can also be various forms of mobile terminals, such as smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.) and other similar mobile terminals.

[0160] Optionally, the electronic device can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 514. The electronic device itself can have a display device, and can also be connected to other display devices through the input / output interface 514. Optionally, a storage device, such as a hard disk, can also be connected through the input / output interface 514, so that data in the electronic device can be stored in the storage device, or data in the storage device can be read, and data in the storage device can also be stored in the memory 511. It can be understood that the input / output interface 514 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 514 can be a component of the electronic device, or it can be an external device connected to the electronic device when needed.

[0161] Optionally, the memory 511 may be a volatile memory and / or a non-volatile memory, the volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory or a flash memory, etc.

[0162] Optionally, the computer program stored in the memory 511 may be divided into one or more modules, which are stored in the memory 511 and executed by the processor 512 to complete the method provided in the embodiment. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device.

[0163] Optionally, the processor 512 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 512 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, and may also be any suitable controller, microcontroller, processor, etc. The processor 512 executes the various methods and processes of this embodiment. Exemplarily, such as a method for training an image processing model or a method for video image processing in an embodiment of this application.

[0164] Optionally, the bus 515 may include a path for transmitting information. The bus 515 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 515 may be divided into an address bus, a data bus, a control bus, etc.

[0165] In an alternative implementation, an embodiment of this application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to execute the methods in the above method embodiments. Part or all of the computer program may be loaded and / or installed on the memory 511 of the electronic device. When the computer program is executed by the processor 512, one or more steps of a method for training an image processing model or a method for video image processing in an embodiment of this application may be executed.

[0166] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0167] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A training method for an image processing model, characterized in that: The method comprises: Acquire several sample images; Performing target object cropping on the sample image to obtain a sample cropped image; Triangulate and deform the sample cropped image to obtain a corresponding deformed sample image; marking the segmented objects on the sample image and the corresponding deformed sample image; The labeled sample cropped image and the corresponding deformed sample image are respectively input into the twin network, and the twin network is trained to obtain a trained twin network.

2. The training method of an image processing model according to claim 1, characterized in that: The marked sample cropped image and the corresponding deformed sample image are respectively input into the twin network, and the twin network is trained to obtain a trained twin network, including: Inputting the marked sample cropped image into the twin network to obtain a first generated image, and inputting the deformed sample image corresponding to the sample cropped image into the twin network to obtain a corresponding second generated image; Obtaining a loss value of the twin network according to the first generated image and the corresponding second generated image; The parameters of the twin network are optimized according to the loss value to obtain a trained twin network.

3. The training method of an image processing model according to claim 2, characterized in that: The obtaining the loss value of the twin network according to the first generated image and the corresponding second generated image includes: performing an inverse deformation operation on the second generated image to obtain an inversely deformed second generated image; A global pixel comparison is performed between the first generated image and the inversely deformed second generated image to obtain a loss value of the twin network.

4. A video image processing method, characterized in that: The method comprises: Extracting video frame images from the video as images to be processed; Cropping the image to be processed according to the region of interest to obtain a cropped image; The cropped image is input into the twin network for processing to obtain a segmented area image; wherein the twin network is a twin network trained using the training method described in any one of claims 1 to 3.

5. A video image processing method according to claim 4, characterized in that: The region of interest is stored in a preset database; Extracting specific location information of the target object in the image to be processed according to the segmented area image; Generate a region of interest of the current video frame image according to the specific location information; The region of interest of the current video frame image is replaced with the region of interest stored in the database and then stored as the region of interest corresponding to the next video frame image.

6. A video image processing method according to claim 4, characterized in that: The method further comprises: According to the segmented area image corresponding to the previous video frame image, a motion vector of each pixel in the segmented area image corresponding to the current video frame image is obtained by a fast optical flow algorithm; According to the motion vector, the noise points or incoherent areas in the segmented area image corresponding to the current video frame image are adjusted to obtain the current segmented area image after post-processing.

7. A video image processing method according to any one of claims 4 to 6, characterized in that: The method further comprises: expanding the boundary of the segmented region image by applying a morphological dilation operation; Performing feathering processing on the boundary to obtain a segmented area image with a faded boundary area; Gaussian blur filtering is performed on the segmented region image with the boundary region faded to obtain a final segmented region image.

8. A video image processing system, characterized in that: The system comprises: An acquisition module is used to extract a video frame image from a video as an image to be processed; A cropping module, used for cropping the image to be processed according to the region of interest to obtain a cropped image; A generation module is used to input the cropped image into a twin network for processing to obtain a segmented area image; wherein the twin network is a twin network trained using the training method described in any one of claims 1 to 3.

9. An electronic device, characterized in that: include: a memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the training method of the image processing model as described in any one of claims 1 to 3 or the video image processing method as described in any one of claims 4 to 7.

10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement a training method for an image processing model as described in any one of claims 1 to 3 or a video image processing method as described in any one of claims 4 to 7 when executed.

Citation Information

Patent Citations

  • Target tracking method based on twin neural network and parallel attention module

    CN111354017A

  • Target tracking method based on conditional adversarial generative twinning network

    CN112837344A

  • Twin network single-target visual tracking method based on difficult sample mining

    CN113888595A

  • Image processing method, image deformation method and fitness evaluation method

    CN115393889A

  • Small sample training method for certificate authentication

    CN117636120A