Video processing method and device, computer program product and electronic equipment
By capturing facial images in response to video stream feature information and generating complete videos when there is a pause, the problem of high resource consumption and insufficient image realism in existing technologies is solved, achieving efficient video processing and natural image continuity.
Patent Information
- Application Number
- CN202511197169.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video processing technologies consume a lot of resources when handling dynamic scenes and have difficulty adapting to changes in the speaker. Video interruption compensation technology also struggles to maintain the naturalness and realism of the image.
By responding to the feature information of the current video stream to trigger acquisition, the face region is located and the face image is stored. When the video stream is stuck, the video is generated and replaced to complete the video, and the face image is used to generate a natural complete video.
It reduces resource consumption, improves speaker recognition accuracy, and ensures the continuity of the video stream and the realism of stuttering.
Smart Images

Figure CN120980294A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer, and in particular, to a video processing method and device, a computer program product and an electronic device. BACKGROUND
[0002] In the field of video processing and enhancement, the existing technology mainly relies on traditional image and video acquisition methods, which face many challenges in processing dynamic scenes and implementing video streams. For example, traditional video acquisition techniques usually require periodic acquisition of video frames, which not only consumes a large amount of resources, but also is difficult to adapt to scenarios where the speaker changes. In addition, when the video stalls during video playback, the current video completion technology often uses simple frame insertion or frame duplication methods, which are difficult to maintain the naturalness and realism of the picture.
[0003] Therefore, there is an urgent need to provide a new video processing method.
[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The purpose of the present disclosure is to provide a video processing method, a video processing device, a computer program product and an electronic device, thereby at least partially overcoming the problem of consuming a large amount of resources in video acquisition and lack of realism in video completion in the related art.
[0006] According to one aspect of the present disclosure, a video processing method is provided, comprising:
[0007] In response to a current video stream, determining first feature information corresponding to the current video stream, and triggering acquisition of the current video stream when the first feature information is determined to be target feature information;
[0008] Locating a face region of the current video stream, obtaining a face image corresponding to the face region, and storing the face image;
[0009] When the current video stream is determined to be stalled, obtaining a first face image corresponding to the current video stream, generating a completed video based on the first face image, and replacing the current video stream with the completed video.
[0010] According to one aspect of the present disclosure, a video processing device is provided, comprising:
[0011] The acquisition triggering module is used to respond to the current video stream, determine the first feature information corresponding to the current video stream, and trigger the acquisition of the current video stream when the first feature information is determined to be the target feature information;
[0012] The video acquisition module is used to locate the face region of the current video stream, acquire the face image corresponding to the face region, and store the face image;
[0013] The video replacement module is used to, when it is determined that the current video stream is stuck, acquire a first face image corresponding to the current video stream, generate a completed video based on the first face image, and replace the current video stream with the completed video.
[0014] According to one aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the video processing method described in any of the exemplary embodiments above.
[0015] According to one aspect of this disclosure, an electronic device is provided, comprising:
[0016] A processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the video processing method according to any of the exemplary embodiments described above by executing the executable instructions.
[0017] This disclosure provides a video processing method that, in response to a current video stream, determines first feature information corresponding to the current video stream; when the first feature information is determined to be target feature information, triggers the acquisition of the current video stream; locates a face region in the current video stream, acquires a face image corresponding to the face region, and stores the face image; when it is determined that the current video stream is choppy, acquires a first face image corresponding to the current video stream, generates a completed video based on the first face image, and uses the completed video to replace the current video stream. On the one hand, during video playback, the system responds to the current video stream, determines the first feature information corresponding to the current video stream, and triggers the acquisition of the current video stream when the first feature information is determined to be the target feature information. This reduces resource consumption and solves the problem of high resource consumption caused by periodically acquiring video frames in related technologies. On the other hand, during video acquisition, the system locates the face region in the current video stream, acquires the face image corresponding to the face region, and stores the face image. This improves the recognition accuracy of the current speaker in the video and solves the problem of difficulty in adapting to changes in the speaker in dynamic scenes in related technologies. Furthermore, when the current video stream is choppy, the system acquires the first face image corresponding to the current video stream and generates a complete video based on the first face image. This reduces computational complexity, ensures the continuity of the images in the current video stream, and, since it acquires the first face image of the person currently speaking in the current video stream and generates a complete video based on the first face image, the completed video image is closer to a natural state, improving the realism of the playback when the video is choppy.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0020] Figure 1 A flowchart illustrating a video processing method according to an example embodiment of the present disclosure is shown schematically.
[0021] Figure 2 The flowchart illustrates a method for locating a face region in a current video stream, obtaining a face image corresponding to the face region, and storing the face image, according to an example embodiment of the present disclosure.
[0022] Figure 3The flowchart illustrates a method for storing a face image into a preset feature library according to an example embodiment of the present disclosure.
[0023] Figure 4 The flowchart illustrates a method for obtaining a first face image corresponding to the current video stream when it is determined that the current video stream is stuttering, according to an example embodiment of the present disclosure.
[0024] Figure 5 The illustration shows a flowchart of a method for generating a completed video based on a first face image and replacing the current video stream using the completed video, according to an example embodiment of the present disclosure.
[0025] Figure 6 A block diagram of a video processing apparatus according to an exemplary embodiment of the present disclosure is shown schematically.
[0026] Figure 7 An electronic device for implementing the video processing method described above is illustrated according to an example embodiment of the present disclosure. Detailed Implementation
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0028] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0029] In the field of video processing and enhancement, existing technologies mainly rely on traditional image and video acquisition methods, which face numerous challenges when handling dynamic scenes and implementing video streams. For example, traditional video acquisition techniques typically require periodically capturing video frames, which not only consumes significant resources but also struggles to adapt to changing speaker scenarios. Furthermore, when video playback experiences stuttering, current video completion techniques often employ simple frame interpolation or frame duplication methods, which fail to maintain the naturalness and realism of the footage.
[0030] To address the aforementioned issues, some explorations have been undertaken in related fields. For example, the paper "Learning Physical-Spatio-Temporal Features for Video Shadow Removal" proposes a dynamic scene shadow removal model, PSTNet, based on physical properties, spatial relationships, and temporal consistency. It estimates local illumination through a physical branch and enhances physical features using a mask-guided attention strategy. However, this method primarily focuses on shadow removal and does not address video completion and speaker recognition. Another paper, "Image-Guided Depth Sampling and Reconstruction," explores image-based adaptive sampling and reconstruction strategies, proposing a superpixel-based sampling and reconstruction algorithm. While it achieves progress in depth map reconstruction, it also fails to address video completion and speaker recognition. None of these solutions systematically solve the core problems in current technologies.
[0031] To address one or more of the aforementioned issues, this exemplary embodiment first provides a video processing method that can be applied to terminal devices. Of course, those skilled in the art can also run the method of this invention on other platforms as needed, and this exemplary embodiment does not impose any special limitations on this. (Reference) Figure 1 As shown, the video processing method may include steps S110-S130:
[0032] Step S110. Respond to the current video stream, determine the first feature information corresponding to the current video stream, and when the first feature information is determined to be target feature information, trigger the acquisition of the current video stream;
[0033] Step S120. Locate the face region of the current video stream, obtain the face image corresponding to the face region, and store the face image;
[0034] Step S130. When it is determined that the current video stream is stuck, a first face image corresponding to the current video stream is obtained, a complete video is generated based on the first face image, and the current video stream is replaced using the complete video.
[0035] The above video processing method responds to the current video stream, determines the first feature information corresponding to the current video stream, and triggers the acquisition of the current video stream when the first feature information is determined to be the target feature information; locates the face region of the current video stream, acquires the face image corresponding to the face region, and stores the face image; when it is determined that the current video stream is stuck, acquires the first face image corresponding to the current video stream, generates a completed video based on the first face image, and uses the completed video to replace the current video stream. On the one hand, during video playback, the system responds to the current video stream, determines the first feature information corresponding to the current video stream, and triggers the acquisition of the current video stream when the first feature information is determined to be the target feature information. This reduces resource consumption and solves the problem of high resource consumption caused by periodically acquiring video frames in related technologies. On the other hand, during video acquisition, the system locates the face region in the current video stream, acquires the face image corresponding to the face region, and stores the face image. This improves the recognition accuracy of the current speaker in the video and solves the problem of difficulty in adapting to changes in the speaker in dynamic scenes in related technologies. Furthermore, when the current video stream is choppy, the system acquires the first face image corresponding to the current video stream and generates a complete video based on the first face image. This reduces computational complexity, ensures the continuity of the images in the current video stream, and, since it acquires the first face image of the person currently speaking in the current video stream and generates a complete video based on the first face image, the completed video image is closer to a natural state, improving the realism of the playback when the video is choppy.
[0036] The following provides a detailed explanation and description of each step involved in the video processing method of the exemplary embodiments of this disclosure.
[0037] In step S110, in response to the current video stream, first feature information corresponding to the current video stream is determined. When the first feature information is determined to be target feature information, the acquisition of the current video stream is triggered.
[0038] The current video stream is the video stream corresponding to the currently playing video. The first feature information of the current video stream may include the voiceprint features, semantic features, emotional features, and speech behavior features of the current video stream. After obtaining the first feature information of the current video stream, it is compared with a preset feature library. When no feature information similar to the first feature information is found, the first feature information is determined to be the target feature information. At this time, the unique current video stream can be triggered for acquisition.
[0039] In one exemplary embodiment, determining the first feature information corresponding to the current video stream in response to the current video stream includes:
[0040] The current video stream is analyzed to extract its voiceprint features, semantic features, emotional features, and speech behavior features.
[0041] Specifically, the current video stream can be analyzed to extract multi-dimensional speech features, which constitute the first feature information. This first feature information includes voiceprint features, semantic features, sentiment features, and speech behavior features. Voiceprint feature extraction employs a combination of MFCC (Mel frequency cepstral coefficients) and i-vector. MFCC obtains the spectral envelope features by converting the speech signal through a Mel filter bank and discrete cosine transform. i-vector maps high-dimensional features to a low-dimensional identity space based on a factor analysis model. Semantic feature extraction utilizes Wav2Vec2.0 to convert audio to text and combines it with TF-IDF (Term Frequency-Inverse Document Frequency) to extract key semantics. Sentiment features are classified using an LSTM (Long Short Term Memory) sentiment analysis network for features such as pitch and speech rate. Speech behavior features are extracted through temporal domain analysis to obtain parameters such as rhythm (syllable interval), pitch (fundamental frequency variation), and pauses (duration of silent segments).
[0042] In one exemplary embodiment, for mobile applications, MobileNetV3 can be used to replace the traditional LSTM sentiment analysis network, and knowledge distillation of Wav2Vec 2.0 can be performed. The teacher model is the original pre-trained model, and the number of parameters in the student model is reduced to one-quarter, while the feature extraction accuracy remains above 92%. After adopting the MobileNetV3 model, the number of parameters is reduced by 60%, and the inference speed is improved by 35%.
[0043] In one exemplary embodiment, multimodal features can be fused, that is, visual features (such as facial micro-expressions and gestures) are introduced into the target feature information for cross-modal joint analysis. Micro-expressions are identified using the FER (Facial Expression Recognition) model, and gestures are detected using OpenPose (an open-source library based on deep learning for real-time estimation of human keypoints). This improves the semantic understanding depth of scene acquisition. By combining features from two different modalities, speech and vision, a more comprehensive understanding of information in the scene can be achieved, and more hidden semantic content can be obtained.
[0044] In one exemplary embodiment, triggering the acquisition of the current video stream when the first feature information is determined to be target feature information includes:
[0045] The voiceprint features, semantic features, and emotional features in the first feature information are compared with a preset feature library. When the matching degree between the voiceprint features and the voiceprint features in the preset feature library is less than a first preset threshold and the preset feature library does not include the semantic features or the emotional features, the first feature information of the current video stream is determined to be the target feature information.
[0046] Trigger image capture of the current video stream.
[0047] Specifically, after extracting the first feature information from the video stream, the first feature information is compared with features in a preset feature library. If the similarity of the voiceprint feature is less than a first preset threshold, and there are no similar semantic or emotional features in the preset feature library, the first feature information is determined to be the target feature information. That is, when it is determined that there are no features similar to the first feature information in the preset feature library, image acquisition of the current video stream is triggered. The first preset threshold can be 0.7, but this example embodiment does not impose a specific limitation on it.
[0048] After fusing multi-dimensional speech features, the data is compared with features stored in a preset feature library. Video capture is triggered only when a new speaker is added, avoiding the waste of resources in traditional periodic capture and improving the accuracy of speaker recognition.
[0049] In step S120, the face region of the current video stream is located, the face image corresponding to the face region is obtained, and the face image is stored.
[0050] When capturing images from the current video stream, the human faces in the current video stream can be captured. That is, the face region in the current video stream is defined, and the face image corresponding to the face region is obtained. After the face image is obtained, it can be stored in a preset feature library.
[0051] In one exemplary embodiment, reference is made to Figure 2 As shown, locating the face region of the current video stream, obtaining the face image corresponding to the face region, and storing the face image include:
[0052] Step S210. Use a face detection algorithm to determine the face region in the current video stream;
[0053] Step S220. Obtain the face image corresponding to the face region according to the image sharpness evaluation method, and store the face image in a preset feature library.
[0054] The following will further explain and illustrate steps S210 and S220. Specifically, a face detection algorithm can be used to locate the face region in the current video stream, and the clearest frontal video frame can be selected as the face image using an image sharpness evaluation method. After obtaining the face image, the face image is stored in a preset feature library. The face detection algorithm can be MTCNN (Multi-task Cascaded Convolutional Networks), and the image sharpness evaluation method can be the gradient variance method. In this example embodiment, no specific limitations are made on the face detection algorithm and the image sharpness evaluation method.
[0055] In one exemplary embodiment, reference is made to Figure 3 As shown, storing the face image into a preset feature library includes:
[0056] Step S310. Normalize the size of the face image to obtain a normalized face image;
[0057] Step S320. Denoise the normalized face image and enhance the local contrast of the normalized face image to obtain the target face image, and store the target face image in a preset feature library.
[0058] The following will further explain and illustrate steps S310 and S320. Specifically, the face image is normalized in size. For example, its size can be normalized to a fixed resolution to unify the input format, resulting in a normalized face image. The fixed resolution can be 256×256, which is not specifically limited in this example embodiment. After obtaining the normalized face image, it is denoised. Denoising can be performed using a bilateral filtering algorithm to suppress Gaussian noise while preserving edge details. The local contrast of the normalized face image can also be enhanced. This can be achieved using CLAHE (Contrast Limiting Adaptive Histogram Equalization) to ensure consistent image brightness under different lighting conditions. After size normalization, denoising, and enhancement of local contrast, the target face image is obtained. After obtaining the target face image, it is stored in a preset feature library.
[0059] In one exemplary embodiment, the method further includes:
[0060] Obtain the background features of the face image;
[0061] The average brightness and contrast of the face image are detected to determine the illumination characteristics;
[0062] Using a face pose assessment model, pose features are determined;
[0063] The background features, lighting features, pose features, face image, and target feature information are associated and stored in a preset feature library.
[0064] Specifically, background features of the current video stream can also be obtained. This can be achieved by separating the foreground and background of the face image using an image segmentation model, thus obtaining background features. Illumination conditions can also be detected on the face image to obtain illumination features, specifically the average brightness and contrast of the face image. After obtaining these features, the features, the face image, and the target feature information can be associated and stored in a preset feature library. The target feature information is the speech features of the current video stream, specifically including voiceprint features, semantic features, emotional features, and speech behavior features.
[0065] Furthermore, the method also includes:
[0066] The background features, illumination features, pose features, the first face image, and the target feature information are stored using a timestamp-aligned database index structure.
[0067] Specifically, when storing background features, lighting features, pose features, the first face image, and target feature information in association, a timestamp-aligned database index structure can be used to facilitate rapid retrieval later.
[0068] In step S130, when it is determined that the current video stream is stuck, a first face image corresponding to the current video stream is obtained, a complete video is generated based on the first face image, and the current video stream is replaced using the complete video.
[0069] During video playback, the current video stream can be monitored. When a stutter is detected, the first face image corresponding to the current video stream is retrieved from the preset feature library. That is, the first face image corresponding to the person speaking in the current video stream is retrieved from the preset feature library. The first face image is used to generate a complete video, and the generated complete video is used to replace the scene in the current video stream.
[0070] In one exemplary embodiment, reference is made to Figure 4 As shown, when it is determined that the current video stream is experiencing stuttering, the first face image corresponding to the current video stream is acquired, including:
[0071] Step S410. Obtain the packet loss rate, stuttering time, or inter-frame pixel difference of the current video stream. When it is determined that the packet loss rate is greater than a preset packet loss rate, the stuttering time is greater than a preset stuttering time, or the inter-frame pixel difference is less than a preset threshold, the current video stream is determined to be stuttering.
[0072] Step S420. Based on the target feature information of the current video stream, determine the first face image corresponding to the current video stream.
[0073] The following will further explain and illustrate steps S410 and S420. Specifically, real-time monitoring of the current video stream can include monitoring the packet loss rate, the stuttering time, and the frame freeze. Monitoring the packet loss rate can be done by statistically analyzing the percentage of data packets lost per unit time using the RTP / RTCP protocol; a packet loss rate exceeding 10% can trigger the generation of a completed video. Monitoring the stuttering time can be done by monitoring the cumulative duration of continuous packet loss; a stuttering time exceeding 500ms can trigger the generation of a completed video. Monitoring the frame freeze can be done by detecting frame freezes through inter-frame pixel differences; when the pixel difference is less than a threshold for more than 5 consecutive frames, a frame freeze is determined, and the generation of a completed video is triggered. When generating the completed video, the target feature information of the current video stream needs to be determined first. Based on the target feature information of the current video stream, a first face image corresponding to the target feature information is obtained from a preset feature library. After obtaining the first face image, a completed video is generated based on the first face image.
[0074] In one exemplary embodiment, determining the first face image corresponding to the current video stream based on the target feature information of the current video stream includes:
[0075] Obtain the voiceprint features of the current video stream, and determine the first face image associated with the voiceprint features in a preset feature library based on the voiceprint features.
[0076] Specifically, the voiceprint features in the target feature information of the current video stream are obtained, and compared with the voiceprint features in a preset feature library to determine the first face image corresponding to the voiceprint features. When obtaining the first face image, the most recently acquired face image can be obtained; that is, the first face image can be determined based on a timestamp.
[0077] In one exemplary embodiment, reference is made to Figure 5 As shown, the step of generating a completed video based on the first face image and replacing the current video stream with the completed video includes:
[0078] Step S510. Determine the phoneme corresponding to the speech in the current video stream, and adjust the lip movements in the first face image according to the phoneme;
[0079] Step S520. Adjust the head movements in the first face image based on affine transformation;
[0080] Step S530. Adjust the eye movements in the first face image;
[0081] Step S540. Based on the adjusted lip movements, adjusted head movements, and adjusted eye movements, generate the completed video and play the completed video.
[0082] The following will further explain and illustrate steps S510-S540. Specifically, the phonemes corresponding to the speech in the current video stream are determined, and the lip movements are adjusted according to the phonemes. The lip movements refer to the degree of mouth opening, which can be identified by Automatic Speech Recognition (ASR) to determine the phonemes corresponding to the current speech. After obtaining the phonemes, the degree of mouth opening is adjusted according to the phonemes. For example, the / a / sound corresponds to a larger degree of mouth opening (the distance between the upper and lower lips accounts for 15%-20% of the height of the face), and the / m / sound corresponds to a closed mouth (the distance is less than 5%). In this example embodiment, the degree of adjustment is not specifically limited.
[0083] The head movements in the first face image can be adjusted using affine transformation. Affine transformation is a linear geometric transformation that preserves the flatness and parallelism of the image. By controlling the number of pixels for horizontal translation of the head in the first face image within a first range and the vertical scaling ratio within a first ratio, natural micro-head movements can be simulated. Specifically, the left-right tilt angle of the head in the first face image can be controlled within a first angle range, and the up-down nodding angle can be controlled within a second angle range. The left-right tilt angle can be ±15°, and the up-down nodding angle can be ±10°.
[0084] The eye movements in the first face image can be adjusted. Specifically, the eye movement model can be combined to simulate opening and blinking. Opening the eyes means the eyelids cover the upper edge of the pupil, and blinking means the eyelids close and open rapidly. The duration of blinking is about 300ms.
[0085] After adjusting the lip movements, eye movements, and head movements in the first face image, a completed video can be generated based on the adjusted lip movements, eye movements, and head movements, and then the completed video can be played.
[0086] In one exemplary embodiment, the method further includes:
[0087] Identify the lip region in the first face image and adjust the lip region.
[0088] Specifically, to reduce the perceived mismatch between the lip shape and the original video, the size of the lip shape area can be adjusted. The lip shape area can be reduced proportionally to 70%-80% of the original size. This example embodiment does not impose a specific limitation on this.
[0089] In one exemplary embodiment, the method further includes:
[0090] The background portion in the first face image is determined, and the background portion is blurred according to the amplitude of the head movement.
[0091] Specifically, the background portion of the first face image can be identified, and then blurred based on the adjusted head movement. Gaussian blur can be used when blurring the background, with a kernel size of 3×3 to 5×5. The blur intensity is positively correlated with the head movement amplitude to enhance the dynamic realism of the image. In this example embodiment, the blur intensity is not specifically limited.
[0092] The generated completed footage replaces the footage lost during the video interruption period. When the packet loss rate is detected to be below 5% and the stuttering time is reset to zero or the footage resumes its dynamic state, the original video stream is seamlessly switched back through timestamp alignment and a fade-in / fade-out transition (lasting 200ms) to ensure viewing continuity.
[0093] In one exemplary embodiment, the parameters can be adjusted according to the computing power of the terminal device. For example, low-performance devices can use a smaller image resolution (128×128) and a simplified bilateral filter kernel (3×3 instead of 5×5), while high-performance devices can retain the original parameters to ensure the real-time performance of different devices, that is, to ensure that the latency is <200ms.
[0094] In this disclosure, when issues such as packet loss, stuttering, or freezing occur in the video stream, complete footage can be generated promptly, giving viewers the impression that the video is uninterrupted and ensuring a continuous viewing experience. Adjustments are made to lip movements based on phonemes, eye movements using an eye movement model, and head movements through affine transformations. Combined with background blurring, complex calculations are avoided, reducing the computational burden on the system. By simulating lip movements, head motion, and background blurring, the completed video footage is made more natural, solving the problem of insufficient realism when video is lost.
[0095] In one exemplary embodiment, the method further includes:
[0096] When the stuttering in the current video stream disappears, the completed video is transitioned by aligning the timestamps to switch back to the current video stream.
[0097] Specifically, when the current video stream stutter disappears, the system can switch back to playing the original video. During the switch, the completed video can be faded in and out using timestamp alignment. For example, a 200ms fade-in / fade-out transition can be used to switch back to the original video, ensuring the continuity of video playback. The system determines that the stuttering in the current video stream has disappeared when the packet loss rate is less than 5% and the stuttering is cleared or the picture is restored.
[0098] In this example embodiment, the lightweight dynamic adjustment strategy based on historical face images significantly reduces computational complexity while maintaining the realism of the image, making it suitable for application scenarios that are sensitive to latency and resource consumption, such as real-time video communication and online education.
[0099] This disclosure also provides a video processing apparatus, with reference to exemplary embodiments thereof. Figure 6 As shown, it may include: a data acquisition triggering module 610, a video acquisition module 620, and a video replacement module 630. Wherein:
[0100] The acquisition triggering module 610 is used to respond to the current video stream, determine the first feature information corresponding to the current video stream, and trigger the acquisition of the current video stream when the first feature information is determined to be the target feature information;
[0101] The video acquisition module 620 is used to locate the face region of the current video stream, acquire the face image corresponding to the face region, and store the face image;
[0102] The video replacement module 630 is used to, when it is determined that the current video stream is stuck, acquire a first face image corresponding to the current video stream, generate a completed video based on the first face image, and replace the current video stream with the completed video.
[0103] The specific details of each module in the aforementioned video processing device have been described in detail in the corresponding video processing methods, and therefore will not be repeated here.
[0104] In one exemplary embodiment of this disclosure, the acquisition triggering module includes:
[0105] The first feature information extraction module is used to analyze the current video stream and extract the voiceprint features, semantic features, emotional features, and speech behavior features of the current video stream.
[0106] In one exemplary embodiment of this disclosure, the acquisition triggering module includes:
[0107] The feature comparison module is used to compare the voiceprint features, semantic features, and emotional features in the first feature information with a preset feature library. When the matching degree between the voiceprint features and the voiceprint features in the preset feature library is less than a first preset threshold and the preset feature library does not include the semantic features or the emotional features, the first feature information of the current video stream is determined to be the target feature information.
[0108] The image acquisition module is used to trigger image acquisition of the current video stream.
[0109] In one exemplary embodiment of this disclosure, the video acquisition module includes:
[0110] The face detection module is used to determine the face regions in the current video stream using a face detection algorithm;
[0111] The face image acquisition module is used to obtain a face image corresponding to the face region according to the image sharpness evaluation method, and store the face image in a preset feature library.
[0112] In one exemplary embodiment of this disclosure, the face image acquisition module includes:
[0113] An image normalization module is used to normalize the size of the face image to obtain a normalized face image;
[0114] The image storage module is used to denoise the normalized face image and enhance the local contrast of the normalized face image to obtain the target face image, and store the target face image in a preset feature library.
[0115] In one exemplary embodiment of this disclosure, the face image acquisition module includes:
[0116] The background feature acquisition module is used to acquire the background features of the face image;
[0117] The illumination feature determination module is used to perform average brightness and contrast detection on the face image to determine the illumination features;
[0118] The pose feature determination module is used to determine pose features using a face pose evaluation model;
[0119] The feature association storage module is used to associate and store the background features, the illumination features, the pose features, the face image, and the target feature information into a preset feature library.
[0120] In one exemplary embodiment of this disclosure, the feature association storage module includes:
[0121] The timestamp-aligned storage module is used to store the background features, the illumination features, the pose features, the first face image, and the target feature information using a timestamp-aligned database index structure.
[0122] In one exemplary embodiment of this disclosure, the video replacement module includes:
[0123] The stuttering determination module is used to obtain the packet loss rate, stuttering time, or inter-frame pixel difference of the current video stream. When it is determined that the packet loss rate is greater than a preset packet loss rate, the stuttering time is greater than a preset stuttering time, or the inter-frame pixel difference is less than a preset threshold, the current video stream is stuttered.
[0124] The face image determination module is used to determine a first face image corresponding to the current video stream based on the target feature information of the current video stream.
[0125] In one exemplary embodiment of this disclosure, the face image determination module includes:
[0126] The face image matching module is used to obtain the voiceprint features of the current video stream and determine the first face image associated with the voiceprint features in a preset feature library based on the voiceprint features.
[0127] In one exemplary embodiment of this disclosure, the video replacement module includes:
[0128] The lip movement adjustment module is used to determine the phonemes corresponding to the speech in the current video stream, and adjust the lip movements in the first face image according to the phonemes;
[0129] A head motion adjustment module is used to adjust the head motion in the first face image based on affine transformation;
[0130] An eye movement adjustment module is used to adjust the eye movements in the first face image;
[0131] The video completion generation module is used to generate the completed video based on the adjusted lip movements, adjusted head movements, and adjusted eye movements, and then play the completed video.
[0132] In one exemplary embodiment of this disclosure, the video replacement module includes:
[0133] The lip region adjustment module is used to determine the lip region in the first face image and adjust the lip region.
[0134] In one exemplary embodiment of this disclosure, the video replacement module includes:
[0135] The background blur module is used to determine the background portion in the first face image and blur the background portion according to the amplitude of the head movement.
[0136] In one exemplary embodiment of this disclosure, the video replacement module includes:
[0137] The original video playback module is used to transition the supplementary video by aligning the timestamps when the stuttering of the current video stream disappears, so as to switch to the current video stream.
[0138] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0139] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0140] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0141] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0142] The following reference Figure 7 To describe an electronic device 700 according to such an embodiment of the present disclosure. Figure 7 The electronic device 700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0143] like Figure 7As shown, the electronic device 700 is manifested in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one processing unit 710, at least one storage unit 720, a bus 730 connecting different system components (including storage unit 720 and processing unit 710), and a display unit 740.
[0144] The storage unit stores program code that can be executed by the processing unit 710, causing the processing unit 710 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 710 can perform actions such as... Figure 1 The steps shown are as follows: S110: Responding to the current video stream, determining the first feature information corresponding to the current video stream; when the first feature information is determined to be target feature information, triggering the acquisition of the current video stream; S120: Locating the face region of the current video stream, acquiring the face image corresponding to the face region, and storing the face image; S130: When it is determined that the current video stream is stuck, acquiring the first face image corresponding to the current video stream, generating a completed video based on the first face image, and using the completed video to replace the current video stream.
[0145] Storage unit 720 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 7201 and / or cache memory 7202, and may further include a read-only memory (ROM) 7203.
[0146] The storage unit 720 may also include a program / utility 7204 having a set (at least one) program module 7205, such program module 7205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0147] Bus 730 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0148] Electronic device 700 can also communicate with one or more external devices 800 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 700, and / or with any device that enables electronic device 700 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 750. Furthermore, electronic device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 760. As shown, network adapter 760 communicates with other modules of electronic device 700 via bus 730. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0149] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0150] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.
[0151] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0152] The program product may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, and random access memory. ( RAM )、 Read-only memory (ROM), erasable programmable read-only memory ( EPROM or flash memory )、 Optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0153] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0154] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0155] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0156] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0157] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A video processing method, characterized in that, include: In response to the current video stream, determine the first feature information corresponding to the current video stream, and when the first feature information is determined to be the target feature information, trigger the acquisition of the current video stream; Locate the face region in the current video stream, obtain the face image corresponding to the face region, and store the face image; When it is determined that the current video stream is stuck, a first face image corresponding to the current video stream is obtained, a complete video is generated based on the first face image, and the complete video is used to replace the current video stream.
2. The method according to claim 1, characterized in that, The step of responding to the current video stream and determining the first feature information corresponding to the current video stream includes: The current video stream is analyzed to extract its voiceprint features, semantic features, emotional features, and speech behavior features.
3. The method according to claim 2, characterized in that, When the first feature information is determined to be target feature information, triggering the acquisition of the current video stream includes: The voiceprint features, semantic features, and emotional features in the first feature information are compared with a preset feature library. When the matching degree between the voiceprint features and the voiceprint features in the preset feature library is less than a first preset threshold and the preset feature library does not include the semantic features or the emotional features, the first feature information of the current video stream is determined to be the target feature information. Trigger image capture of the current video stream.
4. The method according to claim 1, characterized in that, The steps of locating the face region in the current video stream, obtaining the face image corresponding to the face region, and storing the face image include: The face region in the current video stream is determined using a face detection algorithm; A face image corresponding to the face region is obtained according to the image sharpness evaluation method, and the face image is stored in a preset feature library.
5. The method according to claim 4, characterized in that, The step of storing the face image into a preset feature library includes: The face image is normalized to obtain a normalized face image; The normalized face image is denoised and its local contrast is enhanced to obtain the target face image, which is then stored in a preset feature library.
6. The method according to claim 4, characterized in that, The method further includes: Obtain the background features of the face image; The average brightness and contrast of the face image are detected to determine the illumination characteristics; Using a face pose assessment model, pose features are determined; The background features, lighting features, pose features, face image, and target feature information are associated and stored in a preset feature library.
7. The method according to claim 6, characterized in that, The method further includes: The background features, illumination features, pose features, the first face image, and the target feature information are stored using a timestamp-aligned database index structure.
8. The method according to claim 1, characterized in that, When it is determined that the current video stream is experiencing stuttering, the first face image corresponding to the current video stream is acquired, including: The packet loss rate, stuttering time, or inter-frame pixel difference of the current video stream is obtained. When the packet loss rate is greater than a preset packet loss rate, the stuttering time is greater than a preset stuttering time, or the inter-frame pixel difference is less than a preset threshold, the current video stream is determined to be stuttering. Based on the target feature information of the current video stream, a first face image corresponding to the current video stream is determined.
9. The method according to claim 1, characterized in that, The step of determining the first face image corresponding to the current video stream based on the target feature information of the current video stream includes: Obtain the voiceprint features of the current video stream, and determine the first face image associated with the voiceprint features in a preset feature library based on the voiceprint features.
10. The method according to claim 1, characterized in that, The step of generating a completed video based on the first face image and replacing the current video stream with the completed video includes: Determine the phonemes corresponding to the speech in the current video stream, and adjust the lip movements in the first face image according to the phonemes; The head movements in the first face image are adjusted based on affine transformation; Adjust the eye movements in the first face image; The completed video is generated based on the adjusted lip movements, adjusted head movements, and adjusted eye movements, and then played.
11. The method according to claim 10, characterized in that, The method further includes: Identify the lip region in the first face image and adjust the lip region.
12. The method according to claim 10, characterized in that, The method further includes: The background portion in the first face image is determined, and the background portion is blurred according to the amplitude of the head movement.
13. The method according to claim 1, characterized in that, The method further includes: When the stuttering in the current video stream disappears, the completed video is transitioned by aligning the timestamps to switch back to the current video stream.
14. A video processing apparatus, characterized in that, include: The acquisition triggering module is used to respond to the current video stream, determine the first feature information corresponding to the current video stream, and trigger the acquisition of the current video stream when the first feature information is determined to be the target feature information; The video acquisition module is used to locate the face region of the current video stream, acquire the face image corresponding to the face region, and store the face image; The video replacement module is used to, when it is determined that the current video stream is stuck, acquire a first face image corresponding to the current video stream, generate a completed video based on the first face image, and use the completed video to replace the current video stream.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 13.
16. An electronic device, characterized in that, include: Processor; and A memory for storing executable instructions of the processor; wherein the processor is configured to perform the method of any one of claims 1-13 by executing the executable instructions.