Video data processing method, device and system and electronic equipment
By cropping and enhancing facial images in video conferencing images, the image quality degradation problem caused by JPEG compression is solved, achieving efficient image quality improvement and user experience optimization.
Patent Information
- Application Number
- CN202510697497.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
In video conferencing systems, traditional JPEG compression technology leads to image quality degradation, especially block effects and artifacts at high compression ratios, which affects user experience.
By receiving the compressed video image stream, cropping the face image and performing image enhancement, the model is trained using the face image dataset to restore the detailed features, and then the enhanced face image is spliced with the original image to reduce the splicing seams and display the enhanced video image.
While maintaining a high compression ratio, it significantly improves image clarity and naturalness, optimizes user experience, and reduces data transmission bandwidth requirements.
Smart Images

Figure CN120640014A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a video data processing method, device, system and electronic equipment. Background Art
[0002] In video conferencing systems, due to the large amount of image data transmitted and the high bandwidth requirements, simply adding channels is not an effective solution to information transmission and storage issues. Therefore, image compression becomes a key processing method. Traditional JPEG (Joint Photographic Experts Group) compression technology is widely used, but high compression ratios can lead to image quality degradation. In real-time image transmission scenarios such as video conferencing, image clarity and naturalness are crucial to the user experience. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a video data processing method, device, system and electronic equipment, which can improve the quality of the image while ensuring the compression transmission efficiency, and ensure the clarity and fluency of the video conference.
[0004] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0005] In one aspect, a method for processing video data is provided, comprising:
[0006] receiving an image data stream, wherein the image data stream includes a plurality of frames of compressed first video images;
[0007] cropping a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image;
[0008] performing image enhancement on the facial image;
[0009] splicing the enhanced facial image with the second video image to obtain a third video image;
[0010] The third video image is displayed on a display terminal.
[0011] In some embodiments, performing image enhancement on the facial image includes:
[0012] Establishing a facial image dataset, wherein the facial image dataset includes multiple facial images under different lighting conditions, backgrounds, and person postures;
[0013] Using the facial image dataset to train a face enhancement model;
[0014] The trained face enhancement model is used to perform image enhancement on face images.
[0015] In some embodiments, the image data stream is transmitted using the User Datagram Protocol (UDP) and the Real-time Transport Protocol (RTP).
[0016] In some embodiments, cropping a face image from each frame of the first video image includes:
[0017] A face region is identified from the first video image using face recognition, and the face region is cropped from the first video image to obtain the face image.
[0018] In some embodiments, the step of splicing the enhanced facial image with the second video image to obtain a third video image includes:
[0019] The enhanced face image is pasted into the face region of the second video image, and weighted averaging processing is performed on pixel values at the junction of the face image and the second video image.
[0020] An embodiment of the present invention further provides a video data processing device, comprising:
[0021] A receiving module, configured to receive an image data stream, wherein the image data stream includes a first video image after multiple frames of compression;
[0022] a cropping module, configured to crop a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image;
[0023] An image enhancement module, configured to perform image enhancement on the facial image;
[0024] a splicing module, configured to splice the enhanced face image with the second video image to obtain a third video image;
[0025] A display processing module is used to display the third video image on a display terminal.
[0026] An embodiment of the present invention further provides a video data processing system, comprising:
[0027] A video acquisition device is used to acquire video images, compress multiple frames of video images to obtain an image data stream, and send the image data stream to a video data processing device;
[0028] The video data processing device is used to receive an image data stream, which includes multiple frames of compressed first video images; crop a facial image from each frame of the first video image to obtain a facial image and a second video image after removing the facial image; perform image enhancement on the facial image; splice the enhanced facial image with the second video image to obtain a third video image; and display the third video image on a display terminal.
[0029] An embodiment of the present invention further provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program implements the steps of the above-described video data processing method when executed by the processor.
[0030] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video data processing method described above are implemented.
[0031] An embodiment of the present invention further provides a computer program product, comprising computer instructions, which implement the steps of the video data processing method described above when executed by a processor.
[0032] The embodiments of the present invention have the following beneficial effects:
[0033] In the above scheme, an image data stream including multiple frames of a first video image is received, where the first video image is a compressed image. A face image is cropped from each frame of the first video image to obtain a face image and a second video image after removing the face image; the face image is enhanced; the enhanced face image is spliced with the second video image to obtain a third video image; and the third video image is displayed on a display terminal. This can effectively improve the visual effect of the compressed image while significantly reducing the bandwidth required for data transmission. The technical solution of this embodiment can solve the problems of block effects and artifacts caused by traditional JPEG compression, and improve the clarity and naturalness of the compressed image while maintaining a high compression ratio. The technical solution of this embodiment can be applied to real-time image transmission scenarios such as video conferencing, and can significantly improve image quality and optimize user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of a video data processing method according to an embodiment of the present invention;
[0035] Figure 2 A schematic diagram of video data processing according to an embodiment of the present invention;
[0036] Figure 3A schematic diagram of an image being processed by a video processing server according to an embodiment of the present invention;
[0037] Figure 4 Schematic diagram of the structure of a video data processing device according to an embodiment of the present invention;
[0038] Figure 5 Schematic diagram of the composition of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0040] In related technologies, in real-time image transmission scenarios such as video conferencing systems, JPEG technology is used to compress images in order to reduce the amount of transmitted data. However, high compression ratios will lead to a decrease in image quality, and there will also be block effects and artifacts.
[0041] The embodiments of the present invention provide a video data processing method, device, system and electronic device, which can improve image quality while ensuring compression transmission efficiency, thereby ensuring the clarity and fluency of video conferencing.
[0042] An embodiment of the present invention provides a method for processing video data. Figure 1 Shown, including:
[0043] Step 101: receiving an image data stream, wherein the image data stream includes a plurality of compressed frames of a first video image;
[0044] In this embodiment, the video data processing method can be applied to a video processing server, which can receive an image data stream via a network. On the image acquisition side, a video acquisition device, such as a camera, can capture video images, compress the captured raw video images to obtain a first video image, and send an image data stream including multiple frames of the first video image to a video data processing device.
[0045] Specifically, the User Datagram Protocol (UDP) and the Real-time Transport Protocol (RTP) can be used to transmit image data streams. This mode can achieve real-time data transmission by encapsulating UDP data packets with RTP data and controlling data transmission through RTP, thereby ensuring reliable transmission of image data streams.
[0046] Step 102: cropping a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image;
[0047] Specifically, face recognition can be used to identify a face area from the first video image, and the face area can be cropped from the first video image to obtain the face image and the second video image after removing the face image. Face recognition technology can be used to quickly and easily obtain a face image from the first video image.
[0048] Step 103: performing image enhancement on the face image;
[0049] Specifically, a facial image dataset can be established in advance, which includes multiple facial images under different lighting conditions, backgrounds and character postures; the facial image dataset is used to train a face enhancement model; and the trained face enhancement model is used to perform image enhancement on the cropped facial images, so that the detailed features of the facial images can be restored and the image quality can be improved.
[0050] Step 104: splicing the enhanced face image with the second video image to obtain a third video image;
[0051] Specifically, when the enhanced facial image is pasted into the facial area of the second video image, a seam may appear at the junction of the facial image and the second video image. Therefore, the pixel values at the junction of the facial image and the second video image can be weighted averaged to reduce the generation of seams.
[0052] Step 105: Display the third video image on a display terminal.
[0053] In this embodiment, an image data stream including multiple frames of first video images is received, where the first video images are compressed images, a face image is cropped from each frame of the first video image to obtain a face image and a second video image after removing the face image; image enhancement is performed on the face image; the enhanced face image is spliced with the second video image to obtain a third video image; and the third video image is displayed on a display terminal. This can effectively improve the visual effect of the compressed image while significantly reducing the bandwidth required for data transmission.
[0054] This embodiment is not limited to image enhancement of facial images; image enhancement can also be performed on other parts of the first video image to improve image quality. However, the facial image is the most important part of the first video image, and image enhancement of the facial image can significantly improve image quality while also reducing data processing.
[0055] The technical solution of this embodiment can solve the blocking and artifact issues caused by traditional JPEG compression, improving the clarity and naturalness of the compressed image while maintaining a high compression ratio. The technical solution of this embodiment can be applied in real-time image transmission scenarios such as video conferencing, significantly improving image quality and optimizing the user experience.
[0056] The technical solution of this embodiment can be applied to a video conferencing system. The video conferencing system includes a video acquisition device, a video enhancement server, and a display terminal. Figure 2 FIG. 1 is a schematic diagram of a video conferencing system processing video data according to an embodiment of the present invention. Figure 2 As shown, the video capture device captures video conference images through a camera, uses the video conference images in the real-time video stream as input, and performs JPEG or video (AV1) compression encoding on the video conference images to obtain multi-frame compressed video conference images (i.e., the first video image mentioned above). The multi-frame compressed video conference images are sent to the video enhancement server as an image data stream via the network. By compressing the video conference images, the data size can be reduced and transmission is facilitated. Because video conferencing systems require real-time data transmission and have high data reliability requirements, the UDP+RTP transmission mode can be used during transmission to achieve real-time and reliable data transmission.
[0057] During the transmission process, the image quality may be affected by factors such as noise, and the compression of the video conference image will also reduce the image quality. In order to improve the image quality, a video enhancement server is set at the image data receiving end. After receiving the image data stream, the video enhancement server performs the following operations on the image data stream: Figure 3 The processing shown above obtains an enhanced video conference image (ie, the third video image mentioned above); thereafter, the video enhancement server can transmit the enhanced video conference image to the display terminal for display.
[0058] In a video conferencing scenario, facial images have the greatest impact on viewing experience. Therefore, when enhancing video conferencing images, only facial images may be enhanced to reduce the amount of data processing.
[0059] In this embodiment, the video enhancement server can pre-establish an SD model, which is used to process image cropping and enhancement. First, a training dataset is established to train the SD model. When applied to video conferencing scenarios, a dataset containing images of various video conferencing scenes is collected and created. The images in the dataset cover different lighting conditions, backgrounds, and character postures. The creation of the dataset includes image acquisition, annotation, and preprocessing. For example, multiple image frames can be obtained from a multi-camera conference call and finely labeled to create a large-scale video face dataset.
[0060] After creating a large-scale video face dataset, the data collection end can compress the collected video face dataset, such as using JPEG image compression technology combined with Gaussian noise to compress the image, and finally form a compressed video face dataset. The compressed video face dataset is then transmitted to the video enhancement server. By compressing the video face dataset, the amount of transmitted data can be reduced.
[0061] After receiving a compressed video face dataset, the video enhancement server can use the dataset to train a SD model. The input of the SD model can be video face images compressed using JPEG image compression technology and with fewer detailed features. The output of the SD model can be video face images with more detailed features. Through model training, the SD model learns how to enhance the compressed image, restore detailed features, and thus improve image quality. The detailed features of the video face images include, but are not limited to: the location and shape of key points, the shape and proportion of facial organs (such as the thickness of eyebrows, the curvature of the nose bridge, the thickness of lips, and other local morphological features), local texture information, color and reflectivity, dynamic features related to facial expressions (such as the change in glabellar texture when frowning and the degree of upward movement of the mouth corners when smiling), permanent marks (including personalized features such as moles, scars, dimples, and additional elements such as beards and glasses), and image quality-related features (such as the contrast between clarity and blur). The location and shape of key points include the coordinates and relative positional relationships of key points such as the eyes (pupil center, eye corners), nose (nose tip, nostrils), mouth (upper and lower lip edges, mouth corners), and chin contour.
[0062] After the SD model training is completed, the video enhancement server can use the SD model to process the video conference scene image. Figure 3 As shown, after receiving the image data stream, the video enhancement server performs the following processing on each frame of compressed video conference scene image included in the image data stream:
[0063] The faces in the video conferencing scene image are cropped to obtain a face image and a video conferencing scene image after the face image is cropped. To improve the efficiency of face cropping, face recognition technology can be used to perform face recognition on the video conferencing scene image, and then the recognized faces are cropped to form a data set containing only face images.
[0064] A dataset containing only facial images is fed into the trained SD model to generate enhanced facial images. This enhanced facial image is then pasted back into the video conferencing scene image after the facial images have been cropped, resulting in a higher-quality video conferencing image that is then transmitted over the network to a display terminal for display.
[0065] When the enhanced facial image is pasted back onto the video conferencing scene image after the facial image is cropped, a seam may appear at the junction of the facial image and the video conferencing scene image. In order to reduce the appearance of the seam, the pixel values of the overlapping area of the facial image and the video conferencing scene image can be weighted averaged, that is, the final pixel value of each pixel in the overlapping area is the weighted average of the pixel value of the facial image and the pixel value of the video conferencing scene image.
[0066] Through the above processing, this embodiment can achieve real-time enhancement of video images in video conferencing scenarios, significantly improve image quality, and optimize user experience.
[0067] The embodiment of the present invention further provides a video data processing device, such as Figure 4 Shown, including:
[0068] A receiving module 21 is configured to receive an image data stream, wherein the image data stream includes a first video image after multiple frames of compression;
[0069] In this embodiment, the video data processing device can be applied to a video processing server, which can receive an image data stream via a network. On the image acquisition side, a video acquisition device, such as a camera, can capture video images, compress the captured raw video images to obtain a first video image, and send an image data stream including multiple frames of the first video image to the video data processing device.
[0070] Specifically, the User Datagram Protocol (UDP) and the Real-time Transport Protocol (RTP) can be used to transmit image data streams. This mode can achieve real-time data transmission by encapsulating UDP data packets with RTP data and controlling data transmission through RTP, thereby ensuring reliable transmission of image data streams.
[0071] a cropping module 22, configured to crop a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image;
[0072] Specifically, face recognition can be used to identify a face area from the first video image, and the face area can be cropped from the first video image to obtain the face image and the second video image after removing the face image. Face recognition technology can be used to quickly and easily obtain a face image from the first video image.
[0073] An image enhancement module 23, configured to perform image enhancement on the face image;
[0074] Specifically, a facial image dataset can be established in advance, which includes multiple facial images under different lighting conditions, backgrounds and character postures; the facial image dataset is used to train a face enhancement model; and the trained face enhancement model is used to perform image enhancement on the cropped facial images, so that the detailed features of the facial images can be restored and the image quality can be improved.
[0075] a splicing module 24, configured to splice the enhanced face image with the second video image to obtain a third video image;
[0076] The display processing module 25 is configured to display the third video image on a display terminal.
[0077] In this embodiment, the receiving module receives an image data stream including multiple frames of the first video image, the cropping module crops the face image from each frame of the first video image to obtain the face image and the second video image after removing the face image; the image enhancement module performs image enhancement on the face image; the splicing module splices the enhanced face image with the second video image to obtain a third video image; the display processing module displays the third video image on the display terminal, which can effectively improve the visual effect of the compressed image and significantly reduce the bandwidth required for data transmission. The technical solution of this embodiment can solve the block effect and artifact problems caused by traditional JPEG compression, and achieve the goal of improving the clarity and naturalness of the compressed image while maintaining a high compression ratio. The technical solution of this embodiment can be applied to real-time image transmission scenarios such as video conferencing, which can significantly improve image quality and optimize user experience.
[0078] In some embodiments, the image enhancement module 23 is specifically used to establish a facial image dataset, which includes multiple facial images under different lighting conditions, backgrounds and character postures; use the facial image dataset to train a facial enhancement model; and use the trained facial enhancement model to perform image enhancement on facial images.
[0079] When the enhanced facial image is pasted into the facial area of the second video image, a stitching seam may appear at the junction of the facial image and the second video image. In some embodiments, the stitching module 24 is specifically used to paste the enhanced facial image into the facial area of the second video image, and perform weighted averaging processing on the pixel values at the junction of the facial image and the second video image to reduce the generation of stitching seams.
[0080] An embodiment of the present invention further provides a video data processing system, comprising:
[0081] A video acquisition device is used to acquire video images, compress multiple frames of video images to obtain an image data stream, and send the image data stream to a video data processing device;
[0082] The video data processing device is used to receive an image data stream, which includes multiple frames of compressed first video images; crop a facial image from each frame of the first video image to obtain a facial image and a second video image after removing the facial image; perform image enhancement on the facial image; splice the enhanced facial image with the second video image to obtain a third video image; and display the third video image on a display terminal.
[0083] Please refer to Figure 5 An embodiment of the present invention further provides an electronic device 30, including a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31. When the computer program is executed by the processor 31, the various processes of the above-mentioned video data processing method embodiment are implemented, and the same technical effects can be achieved. To avoid repetition, details will not be given here.
[0084] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the above-described video data processing method embodiment and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0085] An embodiment of the present invention further provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of the above-mentioned video data processing method embodiment and can achieve the same technical effect. To avoid repetition, they are not described here.
[0086] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0088] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A video data processing method, characterized in that: include: receiving an image data stream, wherein the image data stream includes a plurality of frames of compressed first video images; cropping a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image; performing image enhancement on the facial image; splicing the enhanced facial image with the second video image to obtain a third video image; The third video image is displayed on a display terminal.
2. The video data processing method according to claim 1, wherein: The performing image enhancement on the face image comprises: Establishing a facial image dataset, wherein the facial image dataset includes multiple facial images under different lighting conditions, backgrounds, and person postures; Using the facial image dataset to train a face enhancement model; The trained face enhancement model is used to perform image enhancement on face images.
3. The video data processing method according to claim 1, wherein: The image data stream is transmitted using the User Datagram Protocol (UDP) and the Real-time Transport Protocol (RTP).
4. The video data processing method according to claim 1, wherein: The step of cropping a face image from each frame of the first video image comprises: A face region is identified from the first video image using face recognition, and the face region is cropped from the first video image to obtain the face image.
5. The video data processing method according to claim 1, wherein: The step of splicing the enhanced facial image with the second video image to obtain a third video image includes: The enhanced face image is pasted into the face region of the second video image, and weighted averaging processing is performed on pixel values at the junction of the face image and the second video image.
6. A video data processing device, characterized in that: include: A receiving module, configured to receive an image data stream, wherein the image data stream includes a first video image after multiple frames of compression; a cropping module, configured to crop a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image; An image enhancement module, configured to perform image enhancement on the facial image; a splicing module, configured to splice the enhanced face image with the second video image to obtain a third video image; A display processing module is used to display the third video image on a display terminal.
7. A video data processing system, characterized in that: include: A video acquisition device is used to acquire video images, compress multiple frames of video images to obtain an image data stream, and send the image data stream to a video data processing device; The video data processing device is configured to receive an image data stream, the image data stream comprising a plurality of frames of compressed first video images; crop a face image from each frame of the first video image to obtain the face image and a second video image after removing the face image; Performing image enhancement on the facial image; splicing the enhanced facial image with the second video image to obtain a third video image; and displaying the third video image on a display terminal.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the video data processing method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the video data processing method according to any one of claims 1 to 5.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the video data processing method according to any one of claims 1 to 5.