Video processing method and device based on enhanced image definition, and electronic device

CN116681614BActive Publication Date: 2026-09-25PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310671937.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2026-09-25
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

相关技术中,由于虚拟人视频生成模型对输入图像和输出图像的尺寸进行限制,导致生成的虚拟人视频中人脸清晰度较低

Benefits of technology

[0048]本申请提出的基于增强图像清晰度的视频处理方法和装置、电子设备,其通过训练得到的目标生成器对原始视频数据中的原始人脸图像进行分辨率增强处理,得到高分辨率的目标人脸图像,并根据目标人脸图像的图像帧顺序对目标人脸图像进行拼接,得到目标视频数据。由此可知,得到的目标视频数据中虚拟人人脸的清晰度大于原始视频数据中虚拟人人脸的清晰度,即本申请实施例提供的方法实现了对视频数据中虚拟人的人脸图像进行清晰度增强。当根据目标视频数据进行金融场景中的虚拟客服、虚拟销售等应用时,能够增强与对象的交互效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116681614B_ABST
    Figure CN116681614B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video processing method and device based on enhanced image definition, and an electronic device, which belong to the technical field of financial technology. The method comprises: performing blur processing on a sample face image to obtain a blurred face image; performing resolution enhancement processing on the blurred face image by using a preset original generator to obtain an initial face image; adjusting parameters of the original generator according to the initial face image and the sample face image to obtain a target generator; performing image frame processing on original video data to obtain original face images and target image frame ordering data; performing resolution enhancement processing on the original face images by using the target generator to obtain target face images; and splicing the target face images according to the target image frame ordering data as the image frame sequence of the target face images to obtain target video data. The embodiments of the present application can enhance the definition of face images in video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology, and in particular to a video processing method and apparatus, and electronic device based on enhancing image clarity. Background Technology

[0002] Currently, in fintech scenarios, virtual human video technology can be used to achieve applications such as virtual customer service and virtual sales. However, these applications have high requirements for the clarity of facial images in the virtual human videos. In related technologies, due to limitations on the size of input and output images in virtual human video generation models, the facial clarity in the generated virtual human videos is relatively low. Therefore, how to enhance the clarity of facial images in videos has become an urgent technical problem to be solved. Summary of the Invention

[0003] The main objective of this application is to propose a video processing method, apparatus, and electronic device based on enhancing image clarity, which aims to improve the clarity of facial images in video data.

[0004] To achieve the above objectives, a first aspect of this application proposes a video processing method based on enhancing image sharpness, the method comprising:

[0005] The acquired sample face images are blurred to obtain blurred face images;

[0006] The blurred face image is enhanced by a preset original generator to obtain an initial face image;

[0007] The parameters of the original generator are adjusted based on the initial face image and the sample face image to obtain the target generator;

[0008] The original video data is processed by image frame segmentation to obtain the original face image and the target image frame sorting data; wherein, the original video data is video data with a virtual human, the original face image is an image with a virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image;

[0009] The target face image is obtained by performing resolution enhancement processing on the original face image through the target generator;

[0010] The target image frame sorting data is used as the image frame order of the target face image. The target face image is then stitched together according to the target image frame sorting data to obtain the target video data.

[0011] In some embodiments, before performing resolution enhancement processing on the original face image using the target generator to obtain the target face image, the method further includes: updating the target generator, specifically including:

[0012] Obtain sample smiley face images;

[0013] The sample smiley face image is blurred to obtain a blurred smiley face image;

[0014] The target generator performs resolution enhancement processing on the blurred smiley face image to obtain an initial smiley face image;

[0015] The target generator is updated based on the initial smiley face image and the sample smiley face image.

[0016] In some embodiments, updating the target generator based on the initial smiley face image and the sample smiley face image includes:

[0017] The initial smiley face image is compared with the sample smiley face image to obtain the smiley face content loss value;

[0018] The initial smiley face image is analyzed by a preset discriminator to obtain smiley face discrimination values;

[0019] The target generator is updated based on the smile content loss value and the smile discrimination value.

[0020] In some embodiments, comparing the initial smiley face image with the sample smiley face image to obtain a smiley face content loss value includes:

[0021] The initial smiley face image is compared pixel by pixel with the sample smiley face image to obtain the smiley face pixel loss value;

[0022] Feature extraction is performed on the initial smiley face image to obtain the first feature value;

[0023] Feature extraction is performed on the sample smiley face image to obtain the second feature value;

[0024] The first feature value and the second feature value are compared to obtain the smile perception loss value.

[0025] The smile content loss value is obtained based on the smile perception loss value and the smile pixel loss value.

[0026] In some embodiments, adjusting the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator includes:

[0027] The initial face image is compared with the sample face image to obtain the face content loss value;

[0028] The initial face image is analyzed by a preset discriminator to obtain a face discrimination value;

[0029] The parameters of the original generator are adjusted based on the face content loss value and the face discrimination value to obtain the target generator.

[0030] In some embodiments, comparing the initial face image with the sample face image to obtain a face content loss value includes:

[0031] The initial face image is compared pixel by pixel with the sample face image to obtain the face pixel loss value;

[0032] Feature extraction is performed on the initial face image to obtain a third feature value;

[0033] Feature extraction is performed on the sample face image to obtain the fourth feature value;

[0034] The third feature value is compared with the fourth feature value to obtain the face perception loss value;

[0035] The face content loss value is obtained based on the face perception loss value and the face pixel loss value.

[0036] In some embodiments, the step of performing image frame segmentation processing on the original video data to obtain the original face image and target image frame sorting data includes:

[0037] The original video data is subjected to image frame segmentation processing to obtain the original image and the original image frame sorting data; wherein, the original image frame sorting data is used to represent the image frame order of the original image;

[0038] Face recognition is performed on the original image to obtain the original face image, and the original image frame sorting data is used as the image frame order of the original face image.

[0039] To achieve the above objectives, a second aspect of this application provides a video processing apparatus based on enhancing image sharpness, the apparatus comprising:

[0040] The blurring module is used to blur the acquired sample face images to obtain blurred face images;

[0041] The first resolution enhancement processing module is used to perform resolution enhancement processing on the blurred face image through a preset original generator to obtain an initial face image;

[0042] The parameter adjustment module is used to adjust the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator;

[0043] The video data processing module is used to perform image frame segmentation processing on the original video data to obtain the original face image and target image frame sorting data; wherein, the original video data is video data with a virtual human, the original face image is an image with a virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image;

[0044] The second resolution enhancement processing module is used to perform resolution enhancement processing on the original face image through the target generator to obtain the target face image;

[0045] The video data generation module is used to use the target image frame sorting data as the image frame order of the target face image, and to stitch the target face image according to the target image frame sorting data to obtain target video data.

[0046] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0047] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0048] The video processing method, apparatus, and electronic device proposed in this application, based on enhanced image clarity, enhances the resolution of original face images in original video data using a trained target generator to obtain high-resolution target face images. These target face images are then stitched together according to their frame order to obtain target video data. Therefore, the clarity of the virtual face in the obtained target video data is greater than that in the original video data. In other words, the method provided in this application enhances the clarity of virtual face images in video data. When used in applications such as virtual customer service and virtual sales in financial scenarios based on the target video data, it can enhance the interaction with the target user. Attached Figure Description

[0049] Figure 1 This is a flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0050] Figure 2This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0051] Figure 3 This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0052] Figure 4 This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0053] Figure 5 This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0054] Figure 6 This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0055] Figure 7 This is another flowchart of a video processing method for enhancing image clarity according to an embodiment of this application;

[0056] Figure 8 This is a schematic diagram of the structure of a video processing device based on enhancing image clarity according to an embodiment of this application;

[0057] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0059] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0061] First, let's analyze some of the terms used in this application:

[0062] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0063] Virtual humans: refers to the representation of a person's geometric and behavioral characteristics in a computer-generated space, i.e., a virtual environment. Virtual humans can be categorized in several ways: First, based on technology, virtual humans can be divided into algorithm-driven and human-driven types. Algorithm-driven types include real-time AI and facial modeling, while human-driven types generate virtual humans based on captured real-person movements. Second, based on visual dimension, virtual humans can be divided into 2D and 3D types. Third, based on structural composition, virtual humans can be divided into digital and holographic types. Digital types can be viewed online, while holographic types can be viewed in person without the naked eye.

[0064] Pixel-wise loss: This calculates the loss between pixels in the prediction image and the target image. Loss functions include Mean Square Error (MSE), Mean Absolute Error (MAE), and Cross-Entropy Loss. Currently, it is mainly used to predict each pair of pixels in the target variable. Because these loss functions evaluate the class prediction for each pixel vector separately and then average them over all pixels, it can be inferred that each pixel in the image has the same learning ability.

[0065] Perceptual loss compares the feature values ​​obtained from convolving the real image with those obtained from convolving the generated image, aiming to approximate high-level information. It's understandable that image feature extraction yields the following types of feature values: First, for one-layer and two-layer network structures, the extracted feature values ​​include low-level features such as image edges, brightness, and color. Second, for three-layer network structures, the extracted feature values ​​include image texture features. Third, for four-layer network structures, the extracted feature values ​​are discriminative features, including features that distinguish image types. Fourth, for five-layer network structures, the extracted feature values ​​are key discriminative features. Therefore, the feature values ​​compared by perceptual loss are of types three through five. In super-resolution tasks, MSE loss leads to a smoother output image, meaning the output image loses details or high-frequency components; therefore, perceptual loss is needed to enhance the details.

[0066] Currently, in fintech scenarios, virtual human video technology can be used to implement applications such as virtual customer service and virtual sales. However, these applications have high requirements for the clarity of facial images in the virtual human videos. In related technologies, due to limitations on the size of input and output images in virtual human video generation models, the facial clarity in the generated virtual human videos is relatively low, affecting the interaction between the virtual customer service representative and the target audience. Therefore, how to enhance the clarity of facial images in videos has become an urgent technical problem to be solved.

[0067] Based on this, embodiments of this application provide a video processing method, apparatus, and electronic device based on enhanced image clarity, aiming to improve the facial clarity of virtual humans in enhanced video data, thereby improving the interactive effect of virtual customer service and virtual sales.

[0068] The video processing method, apparatus, and electronic device based on enhanced image clarity provided in this application are specifically described through the following embodiments. First, the video processing method based on enhanced image clarity in the embodiments of this application are described.

[0069] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0070] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0071] The video processing method based on enhanced image clarity provided in this application relates to the field of financial technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video processing method based on enhanced image clarity, but is not limited to the above forms.

[0072] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0073] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user image data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments obtained.

[0074] Figure 1 This is an optional flowchart of a video processing method based on enhancing image sharpness provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0075] Step S101: Blur the acquired sample face image to obtain a blurred face image;

[0076] Step S102: The blurred face image is processed by a preset original generator to enhance its resolution and obtain an initial face image.

[0077] Step S103: Adjust the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator;

[0078] Step S104: Perform image frame segmentation processing on the original video data to obtain the original face image and target image frame sorting data; wherein, the original video data is video data with virtual human, the original face image is an image with virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image.

[0079] Step S105: The original face image is enhanced in resolution using the target generator to obtain the target face image;

[0080] Step S106: Use the target image frame sorting data as the image frame order of the target face image, and stitch the target face image according to the target image frame sorting data to obtain the target video data.

[0081] Steps S101 to S106, as illustrated in this embodiment, involve using a trained target generator to enhance the resolution of the original face image in the original video data, obtaining a high-resolution target face image. The target face image is then stitched together according to its frame order to obtain the target video data. Therefore, the clarity of the virtual face in the obtained target video data is greater than that in the original video data. This means the method provided in this embodiment enhances the clarity of the virtual face image in the video data. When used in applications such as virtual customer service and virtual sales in financial scenarios based on the target video data, it can enhance the interaction with the target user.

[0082] In step S101 of some embodiments, multiple sample face images are obtained through a preset face dataset. These sample face images are then blurred to reduce their resolution, resulting in a blurred face image. Specifically, this can be achieved by first convolving the sample face image with a blurring convolution kernel to blur the entire image, and then adding noise. It is understood that the above-described method for obtaining a blurred face image is merely exemplary and is not specifically limited in this application. It is understood that the sample face image is used to train the original generator to obtain a target generator capable of generating clear images. Therefore, to improve the adaptability of the target generator to virtual human video data, the sample face image can also be an image extracted from sample video data. This sample video data consists of video data with virtual humans and clear resolution. The video data can correspond to virtual customer service video data or virtual sales video data in financial scenarios, such as virtual sales video data introducing insurance products, or virtual couple video data used to interact with and answer insurance questions. The corresponding sample face images are images with virtual human faces. It is understood that the sample images can be in any format such as JPEG, TIFF, RAW, GIF, etc., and this application embodiment does not specifically limit this.

[0083] In step S102 of some embodiments, a raw generator based on a Residual in Residual Dense Block (RRDB) structure is pre-set. The RRDB structure not only avoids performance degradation caused by an overly deep generation model, but also improves the representation ability of features. The blurred face image is used as input data to the raw generator to generate a high-resolution image, i.e., the initial face image, through mapping.

[0084] In step S103 of some embodiments, a loss value is calculated for the initial face image and the sample face image according to a preset loss function. The parameters of the original generator are adjusted according to the loss value so that the initial face image generated by the original generator approximates the sample face image. That is, the ability of the original generator to map a low-resolution input image to a high-resolution output image is improved, thereby obtaining the target generator.

[0085] Reference Figure 2 In some implementations, step S103 includes, but is not limited to, steps S201 to S203.

[0086] Step S201: Compare the initial face image with the sample face image to obtain the face content loss value;

[0087] Step S202: The initial face image is discriminated using a preset discriminator to obtain face discrimination values;

[0088] Step S203: Adjust the parameters of the original generator based on the face content loss value and the face discrimination value to obtain the target generator.

[0089] In step S201 of some embodiments, the initial face image and the sample face image are compared at the pixel level, that is, the content, global structure and other high-level information of the initial face image and the sample face image are compared to obtain the face content loss value.

[0090] Reference Figure 3 In some embodiments, step S201 includes, but is not limited to, steps S301 to S305.

[0091] Step S301: Compare the initial face image with the sample face image pixel by pixel to obtain the face pixel loss value;

[0092] Step S302: Extract features from the initial face image to obtain the third feature value;

[0093] Step S303: Extract features from the sample face image to obtain the fourth feature value;

[0094] Step S304: Compare the third feature value with the fourth feature value to obtain the face perception loss value;

[0095] Step S305: Obtain the face content loss value based on the face perception loss value and the face pixel loss value.

[0096] In step S301 of some embodiments, the loss of each pixel vector in the initial face image and the sample face image is calculated by using a preset MSE loss function, MAE loss function, cross-entropy loss, etc., to obtain the face pixel loss value.

[0097] In step S302 of some embodiments, features are extracted from the initial face image using a preset convolutional network model (Visual Geometry Group, VGG) such as VGG-16 or VGG-19 to obtain a third feature value. This third feature value represents high-level features of the initial face image, such as texture features, distinctive features, and key features with discriminative power.

[0098] In step S303 of some embodiments, feature extraction is performed on the sample face image using a preset method such as VGG-16 or VGG-19 to obtain a fourth feature value. The second feature value represents high-level features of the sample face image, such as texture features, distinctive features, and key features with discriminative power. It is understood that the extraction of both the third and fourth feature values ​​can be performed using VGG-16; or both using VGG-19; or the third feature value can be extracted using VGG-16 and the fourth feature value using VGG-19; or the fourth feature value can be extracted using VGG-16 and the third feature value using VGG-19.

[0099] In step S304 of some embodiments, loss calculation is performed on the third feature value and the fourth feature value according to a preset perceptual loss function to obtain the face perception loss value used to characterize the loss of high-level features of the image.

[0100] In step S305 of some embodiments, the face perception loss value and the face pixel loss value are summed to obtain the face content loss value.

[0101] This application embodiment uses the sum of pixel loss and perceptual loss as content loss, which not only realizes the content comparison between the initial face image and the sample face image at the pixel level, but also realizes the comparison of detailed content and global image structure. Therefore, it can improve the clarity of the image generated by the target generator to a certain extent.

[0102] In step S202 of some embodiments, the initial face image is verified for authenticity using a preset discriminator, i.e., it is determined whether the initial face image is a real image or an image generated by the generator, thus obtaining a face discrimination value. It is understood that the discriminator's model structure is based on the U-net network structure. The discriminator's task is to distinguish the high-resolution initial face image generated by the original generator from the label image, i.e., from the sample face image, thereby enabling the high-resolution image generated by the original generator to restore more image texture details and improve the clarity of the image generated by the original generator.

[0103] In step S203 of some embodiments, the original generator is adjusted based on the face content loss value and the face discrimination value to obtain the target generator.

[0104] In step S104 of some embodiments, the raw video data to be processed is obtained through an Application Programming Interface (API) or similar means. This raw video data includes virtual customer service video data and virtual sales video data featuring a virtual human. It is understood that this raw video data can be video data directly generated by a relevant application, such as custom customer service script audio data or sales script audio data input into the relevant application, as well as object video data of the real object that the virtual human is expected to simulate. The generative model in the relevant application is trained based on this audio data and object video data to obtain the corresponding raw video data. It is understood that when this application is applied to insurance scenarios in fintech, customer service scripts include greetings, answers to insurance questions, etc. Sales scripts can be sales pitches for a specific insurance product, such as sales pitches for life insurance or accident insurance. Alternatively, the raw video data can be self-trained video data, i.e., video data generated through self-training steps such as modeling, facial expression design, motion design, and image output. The original video data is processed by image frame segmentation, which involves extracting video frames one by one from the original video data to obtain the original face image and the image frame order of the original face image in the original video data, i.e., the target image frame sorting data.

[0105] It is understood that image frame segmentation processing can be performed using Fast Forward MPEG (FFmpeg) or other methods, and this application embodiment does not specifically limit this. Specifically, the original face image can be obtained by setting the frame rate or the cropping interval. Setting the frame rate sets the number of screenshots per second; for example, when the frame rate is set to 25, 25 images will be extracted from the original video data per second. The cropping interval sets the start time, frame rate, and duration of the cropping; for example, it can be set to start cropping from the 10th second of the original video data, performing the cropping operation at a rate of 25 images per second for a total of 5 seconds. It is understood that this application does not specifically limit the frame rate and cropping interval, but to ensure the clarity of the face in the target video data, image frame segmentation processing should be performed on at least all video segments containing the virtual face in the original video data. Furthermore, the original face image can be in any format such as JPEG, TIFF, RAW, or GIF, and this application embodiment does not specifically limit this.

[0106] Reference Figure 4In some instances, step S104 includes, but is not limited to, steps S401 to S402.

[0107] Step S401: Perform image frame segmentation processing on the original video data to obtain the original image and the original image frame sorting data; wherein, the original image frame sorting data is used to represent the image frame order of the original image;

[0108] Step S402: Perform face recognition on the original image to obtain the original face image, and use the original image frame sorting data as the image frame order of the original face image.

[0109] In step S401 of some embodiments, the original video data is processed by image frame segmentation using ffmpeg or other methods to obtain an original image including a face region and a background region, as well as the original image frame order of the original image in the original video data; or, an invalid image that does not contain a face image and an original image that contains a face image are obtained. That is, image frame segmentation can be performed on video segments containing only face images in the original video data to obtain the original image; or, image frame segmentation can be performed on all video segments of the original video data to obtain an invalid image and the original image.

[0110] In step S402 of some embodiments, when all images obtained from image framing are images containing virtual faces (i.e., when image framing is only performed on video segments containing virtual faces in the original video data), a preset face recognition model is used to perform face recognition processing on the original images to determine the face region representing the virtual face in the original images. This region is then cropped to obtain the original face image, thereby filtering out interference data in the sample images, i.e., filtering out the background region. Alternatively, when the images obtained from image framing include images with virtual faces and images without virtual faces (i.e., when image framing is performed on video segments containing virtual faces in the original video data as well as on video segments without virtual faces in the original video data), a preset face recognition model is used to perform face recognition on invalid images and original images to distinguish between them. Then, the filtered original images are subjected to face recognition again to obtain the original face image. It is understandable that, since the original face image is a partial image of the original image, the original face image and the original image have the same image frame order in the original video data. That is, the original image frame order can be used as the image frame order of the original face image.

[0111] Reference Figure 5 In some embodiments, before step S105, the method provided in this application further includes updating the target generator, specifically including but not limited to steps S501 to S504.

[0112] Step S501: Obtain sample smiley face images;

[0113] Step S502: Blur the sample smiley face image to obtain a blurred smiley face image;

[0114] Step S503: The blurred smiley face image is enhanced by the target generator to obtain the initial smiley face image;

[0115] Step S504: Update the target generator based on the initial smiley face image and the sample smiley face image.

[0116] In step S501 of some embodiments, since virtual humans are often in a speaking state in video data, i.e., they will show their teeth and smile, it is necessary to optimize the target generator using a smiley face dataset to improve the clarity of details such as lips and teeth in the images generated by the target generator. Specifically, multiple sample smiley face images are obtained from a preset smiley face dataset, wherein the sample smiley face images are images with human faces, and the faces are showing their teeth and smiling.

[0117] In step S502 of some embodiments, the sample smiley face image is blurred to reduce its resolution, resulting in a blurred smiley face image. Specifically, this can be achieved by first convolving the sample smiley face image with a blurring convolution kernel to blur the entire image, and then adding noise. It is understood that the above-described method for obtaining the blurred smiley face image is merely exemplary, and this application embodiment does not impose specific limitations on it. Furthermore, the blurring convolution kernel and noise used to obtain the blurred smiley face image can be the same as or different from those used to obtain the blurred face image, and this application embodiment also does not impose specific limitations on them.

[0118] It is understood that the sample smiley face image is used to optimize the target generator, thereby improving the clarity of the lips and teeth in the images generated by the target generator. Therefore, to improve the compatibility between the optimized target generator and the virtual human video data, the sample smiley face image can also be an image extracted from sample video data, which is video data featuring a virtual human and having clear clarity. Correspondingly, the sample smiley face image is an image featuring a virtual human face in a smiling state. It is understood that the sample smiley face image can be in any format such as JPEG, TIFF, RAW, GIF, etc., and this embodiment of the application does not specifically limit this.

[0119] In step S503 of some embodiments, the blurred smiley face image is used as input data to the target generator so that the blurred smiley face image is mapped to a high-resolution image, i.e., the initial smiley face image, by the target generator.

[0120] In step S504 of some embodiments, a loss value is calculated on the initial smiley face image and the sample smiley face image using a preset loss function. The parameters of the target generator are then adjusted based on this loss value, thereby optimizing and updating the target generator.

[0121] Understandably, depending on actual needs, the target generator can also be optimized and trained using training sets of other object features, such as training sets of images corresponding to eyes, hair, ears, and facial contours.

[0122] This application embodiment optimizes and trains the target generator using open-source or preset sample smiley face images to improve the clarity of lips and teeth in the images generated by the target generator, thereby improving the local clarity of the generated images.

[0123] Reference Figure 6 In some embodiments, step S504 may include, but is not limited to, steps S601 to S603.

[0124] Step S601: Compare the image content of the initial smiley face image with that of the sample smiley face image to obtain the smiley face content loss value;

[0125] Step S602: The initial smiley face image is judged by a preset discriminator to obtain the smiley face judgment value;

[0126] Step S603: Update the target generator based on the smile content loss value and the smile discrimination value.

[0127] In step S601 of some embodiments, the initial smiley face image and the sample smiley face image are compared at the pixel level, that is, the content, global structure and other high-level information of the initial smiley face image and the sample smiley face image are compared to obtain the smiley face content loss value.

[0128] Reference Figure 7 In some embodiments, step S601 includes, but is not limited to, steps S701 to S705.

[0129] Step S701: Compare the initial smiley face image with the sample smiley face image pixel by pixel to obtain the smiley face pixel loss value;

[0130] Step S702: Extract features from the initial smiley face image to obtain the first feature value;

[0131] Step S703: Extract features from the sample smiley face image to obtain the second feature value;

[0132] Step S704: Compare the first feature value with the second feature value to obtain the smile perception loss value;

[0133] Step S705: Obtain the smile content loss value based on the smile perception loss value and the smile pixel loss value.

[0134] In step S701 of some embodiments, the loss of each pixel vector in the initial smiley face image and the sample smiley face image is calculated by using a preset MSE loss function, MAE loss function, cross-entropy loss, etc., to obtain the smiley face pixel loss value.

[0135] In step S702 of some embodiments, features are extracted from the initial smiley face image using a preset method such as VGG-16 or VGG-19 to obtain a first feature value. The first feature value represents high-level features of the initial smiley face image, such as texture features, distinctive features, and key features with discriminative power.

[0136] In step S703 of some embodiments, features are extracted from the sample smiley face image using a preset method such as VGG-16 or VGG-19 to obtain a second feature value. The second feature value represents high-level features of the sample smiley face image, such as texture features, distinctive features, and key features with discriminative power. It is understood that the extraction of both the first and second feature values ​​can be performed using VGG-16; or both can be extracted using VGG-19; or the first feature value can be extracted using VGG-16 and the second feature value using VGG-19; or the second feature value can be extracted using VGG-16 and the first feature value using VGG-19.

[0137] In step S704 of some embodiments, the first feature value and the second feature value are calculated according to a preset perceptual loss function to obtain a smile perceptual loss value used to characterize the loss of high-level features of the image.

[0138] In step S705 of some embodiments, the smile perception loss value and the smile pixel loss value are summed to obtain the smile content loss value.

[0139] This application embodiment uses the sum of pixel loss and perceptual loss as content loss, which not only realizes the comparison of the initial smiley face image and the sample smiley face image at the pixel level, but also realizes the comparison of detailed content and global image structure. Therefore, it can improve the clarity of the image generated by the optimized target generator to a certain extent.

[0140] In step S602 of some embodiments, the initial smiley face image is verified for authenticity using a preset discriminator, i.e., it is determined whether the initial smiley face image is a real image or an image generated by the generator, thus obtaining a smiley face discrimination value. It is understood that the discriminator's model structure is built based on the U-net network structure. The discriminator's task is to distinguish the high-resolution initial smiley face image generated by the target generator from the label image, i.e., from the sample smiley face image, thereby enabling the high-resolution image generated by the target generator to restore more image texture details and improve the clarity of the image generated by the target generator.

[0141] In step S603 of some embodiments, the target generator is adjusted based on the smile content loss value and the smile discrimination value, thereby achieving optimization and updating of the target generator.

[0142] Understandably, the unoptimized target generator and the pre-set discriminator constitute a Generative Adversarial Network (GAN). The pre-set discriminator is a pre-trained discriminator, meaning it already possesses the ability to distinguish between real images and generator-generated images. Specifically, the pre-set discriminator can be trained as follows: First, a real image is used as input data for the discriminator, labeled 1, indicating that the current input data is a real image. The discriminator is trained based on its output data and label. Second, an image generated by the pre-set generator is used as input data for the discriminator, labeled 0, indicating that the current input data is a pseudo-image. The discriminator is trained based on its output data and label. These two training steps are repeated until the discriminator converges.

[0143] In step S105 of some embodiments, the original face image is used as input data to the target generator to generate a target face image, wherein the resolution of the target face image is greater than the resolution of the original face image.

[0144] In step S106 of some embodiments, the image frame sorting data of the original face image is used as the image frame order of the target face image in the original video data. This allows the corresponding original face image in the original video data to be replaced with the target face image based on the target image frame sorting data, and then spliced ​​with the preceding and following frames to obtain the target video data. For example, with a frame rate of 1, i.e., 1 second... Figure 1Taking a 20-second original video clip as an example, the clip from second 6 to second 10 contains a virtual human. This clip is processed by frame segmentation to obtain five original face images: A1, A2, A3, A4, and A5. The target image frame order for A1 is second 6, for A2 second 7, for A3 second 8, for A4 second 9, and for A5 second 10. These original face images A1 through A5 are then input into a target generator in batches to obtain target face images B1, B2, B3, B4, and B5. The five target face images are stitched together according to the target image frame sorting data corresponding to each target face image, that is, target face images B1 to B5 are stitched together sequentially to obtain an image set. This image set is then stitched together with video segments from 0 to 5 seconds and from 11 to 20 seconds in the original video data to obtain the target video data.

[0145] It is understood that, in the above example, although the number of seconds is used as the target frame sorting data of the original face image, it should be understood that, depending on actual needs, other data that can describe the image frame order of the original face image in the original video data can also be used as the target image frame sorting data. This application embodiment does not specifically limit this.

[0146] This application embodiment uses a trained target generator to perform resolution enhancement processing on the original face images in the original video data, obtaining high-resolution target face images. The target face images are then stitched together according to their frame order to obtain the target video data. Therefore, the clarity of the virtual face in the obtained target video data is greater than that in the original video data. In other words, the method provided in this application embodiment enhances the clarity of the virtual face images in the video data. When this application is applied to insurance scenarios in fintech, the original video data is either virtual customer service video data featuring virtual humans for the insurance field, or virtual sales video data featuring virtual humans for promoting insurance products. The method described in the above embodiments can enhance the resolution of the virtual face images in the original video data, thereby improving the interaction with insurance consultants and insurance purchasers.

[0147] Please see Figure 8This application also provides a video processing apparatus based on enhanced image clarity, which can implement the above-mentioned video processing method based on enhanced image clarity. The apparatus includes:

[0148] The blurring module 801 is used to blur the acquired sample face image to obtain a blurred face image;

[0149] The first resolution enhancement processing module 802 is used to perform resolution enhancement processing on the blurred face image through a preset original generator to obtain an initial face image;

[0150] The parameter adjustment module 803 is used to adjust the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator;

[0151] The video data processing module 804 is used to perform image frame segmentation processing on the original video data to obtain the original face image and the target image frame sorting data; wherein, the original video data is video data with a virtual human, the original face image is an image with a virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image.

[0152] The second resolution enhancement processing module 805 is used to perform resolution enhancement processing on the original face image through the target generator to obtain the target face image;

[0153] The video data generation module 806 is used to take the target image frame sorting data as the image frame order of the target face image, and stitch the target face image according to the target image frame sorting data to obtain the target video data.

[0154] The specific implementation of this video processing device based on enhanced image clarity is basically the same as the specific embodiment of the video processing method based on enhanced image clarity described above, and will not be repeated here.

[0155] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned video processing method based on enhanced image clarity. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0156] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0157] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0158] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and called and executed by the processor 901 using the video processing method based on enhanced image clarity according to the embodiments of this application.

[0159] The input / output interface 903 is used to implement information input and output;

[0160] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0161] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0162] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0163] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video processing method based on enhanced image clarity.

[0164] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0165] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0166] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0168] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0169] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0170] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0171] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0172] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0173] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0174] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0175] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A video processing method based on enhancing image sharpness, characterized in that, The method includes: The acquired sample face images are blurred to obtain blurred face images; The blurred face image is enhanced by a preset original generator to obtain an initial face image; The parameters of the original generator are adjusted based on the initial face image and the sample face image to obtain the target generator; The original video data is processed by image frame segmentation to obtain the original face image and the target image frame sorting data; wherein, the original video data is video data with a virtual human, the original face image is an image with a virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image; The target face image is obtained by performing resolution enhancement processing on the original face image through the target generator; The target image frame sorting data is used as the image frame order of the target face image. The target face image is then stitched together according to the target image frame sorting data to obtain the target video data. Before performing resolution enhancement processing on the original face image using the target generator to obtain the target face image, the method further includes: updating the target generator, specifically including: Obtain sample smiling face images; wherein, the sample smiling face images are images extracted from sample video data, the sample video data are video data with virtual human and clear resolution, and the sample smiling face images are images with virtual human faces and the faces are in a smiling state. The sample smiley face image is blurred to obtain a blurred smiley face image; The target generator performs resolution enhancement processing on the blurred smiley face image to obtain an initial smiley face image; The target generator is updated based on the initial smiley face image and the sample smiley face image; The step of updating the target generator based on the initial smiley face image and the sample smiley face image includes: The initial smiley face image is compared pixel by pixel with the sample smiley face image to obtain a smiley face pixel loss value; features are extracted from the initial smiley face image to obtain a first feature value; features are extracted from the sample smiley face image to obtain a second feature value; the first feature value and the second feature value are compared to obtain a smiley face perception loss value; and a smiley face content loss value is obtained based on the smiley face perception loss value and the smiley face pixel loss value. The initial smiley face image is discriminated by a preset discriminator to obtain a smiley face discrimination value. The discriminator's task is to distinguish the high-resolution initial smiley face image generated by the target generator from the sample smiley face image, so that the high-resolution image generated by the target generator can restore more image texture details and improve the clarity of the image generated by the target generator. The target generator is updated based on the smile content loss value and the smile discrimination value.

2. The method according to claim 1, characterized in that, The step of adjusting the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator includes: The initial face image is compared with the sample face image to obtain the face content loss value; The initial face image is analyzed by a preset discriminator to obtain a face discrimination value; The parameters of the original generator are adjusted based on the face content loss value and the face discrimination value to obtain the target generator.

3. The method according to claim 2, characterized in that, The step of comparing the initial face image with the sample face image to obtain a face content loss value includes: The initial face image is compared pixel by pixel with the sample face image to obtain the face pixel loss value; Feature extraction is performed on the initial face image to obtain a third feature value; Feature extraction is performed on the sample face image to obtain the fourth feature value; The third feature value is compared with the fourth feature value to obtain the face perception loss value; The face content loss value is obtained based on the face perception loss value and the face pixel loss value.

4. The method according to any one of claims 1 to 3, characterized in that, The step of performing image frame segmentation processing on the original video data to obtain the original face image and target image frame sorting data includes: The original video data is subjected to image frame segmentation processing to obtain the original image and the original image frame sorting data; wherein, the original image frame sorting data is used to represent the image frame order of the original image; Face recognition is performed on the original image to obtain the original face image, and the original image frame sorting data is used as the image frame order of the original face image.

5. A video processing apparatus based on enhancing image clarity, characterized in that, The device includes: The blurring module is used to blur the acquired sample face images to obtain blurred face images; The first resolution enhancement processing module is used to perform resolution enhancement processing on the blurred face image through a preset original generator to obtain an initial face image; The parameter adjustment module is used to adjust the parameters of the original generator based on the initial face image and the sample face image to obtain the target generator; The video data processing module is used to perform image frame segmentation processing on the original video data to obtain the original face image and target image frame sorting data; wherein, the original video data is video data with a virtual human, the original face image is an image with a virtual human face, and the target image frame sorting data is used to represent the image frame order of the original face image; The second resolution enhancement processing module is used to perform resolution enhancement processing on the original face image through the target generator to obtain the target face image; The video data generation module is used to use the target image frame sorting data as the image frame order of the target face image, and to stitch the target face image according to the target image frame sorting data to obtain target video data; Before performing resolution enhancement processing on the original face image through the target generator to obtain the target face image, the device is further configured to: update the target generator, specifically including: Obtain sample smiling face images; wherein, the sample smiling face images are images extracted from sample video data, the sample video data are video data with virtual human and clear resolution, and the sample smiling face images are images with virtual human faces and the faces are in a smiling state. The sample smiley face image is blurred to obtain a blurred smiley face image; The target generator performs resolution enhancement processing on the blurred smiley face image to obtain an initial smiley face image; The target generator is updated based on the initial smiley face image and the sample smiley face image; The step of updating the target generator based on the initial smiley face image and the sample smiley face image includes: The initial smiley face image is compared pixel by pixel with the sample smiley face image to obtain a smiley face pixel loss value; features are extracted from the initial smiley face image to obtain a first feature value; features are extracted from the sample smiley face image to obtain a second feature value; the first feature value and the second feature value are compared to obtain a smiley face perception loss value; and a smiley face content loss value is obtained based on the smiley face perception loss value and the smiley face pixel loss value. The initial smiley face image is discriminated by a preset discriminator to obtain a smiley face discrimination value. The discriminator's task is to distinguish the high-resolution initial smiley face image generated by the target generator from the sample smiley face image, so that the high-resolution image generated by the target generator can restore more image texture details and improve the clarity of the image generated by the target generator. The target generator is updated based on the smile content loss value and the smile discrimination value.

6. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Deblurred face recognition method and system and inspection robot

    CN111460939A

  • Video deblurring method and device, equipment and storage medium

    CN114359079A