Video parameter adjustment method and apparatus
By monitoring and adjusting the quality indicators of digital human live-streaming videos and using neural network models to automatically adjust parameters, the problem of the inability to effectively monitor and adjust the quality of digital human live-streaming videos in existing technologies has been solved, thus improving video quality and live-streaming effects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
- Filing Date
- 2025-07-29
- Publication Date
- 2026-05-07
AI Technical Summary
Existing live video quality monitoring technologies cannot effectively monitor and adjust the advanced semantic information in digital human live videos, resulting in fluctuations in quality such as image clarity, audio-visual synchronization, digital human lip-reading accuracy, and image consistency, which affects the live broadcast effect.
By monitoring quality indicators related to digital humans in live video, such as lip-sync accuracy, audio-visual synchronization, and image consistency, neural network models are used to automatically adjust the parameters of the generated digital humans, thereby achieving feedback adjustment of the live video.
It improves the quality of digital human live streaming videos, ensuring image clarity, audio-visual synchronization, and consistency of the digital human's appearance, thereby enhancing the live streaming experience.
Smart Images

Figure CN2025111189_07052026_PF_FP_ABST
Abstract
Description
Methods and apparatus for adjusting video parameters
[0001] This application claims priority to Chinese Patent Application No. 202411526518.9, filed on October 29, 2024, entitled "Method and Apparatus for Adjusting Video Parameters", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence, and more specifically, to a method and apparatus for adjusting video parameters. Background Technology
[0003] With the development of artificial intelligence technology, digitally generated avatars (digital humans) can replace live streamers in video broadcasts, enabling uninterrupted broadcasts and improving efficiency. However, during a live stream, factors such as server computing power, network fluctuations, and weak network conditions can affect video quality, including but not limited to image clarity, audio-visual synchronization, lip-sync accuracy, and image consistency, thus impacting the broadcast's effectiveness.
[0004] Traditional live video quality monitoring and adjustment technologies primarily target video streaming, monitoring and adjusting streaming parameters such as bitrate stability and image clarity. However, they cannot provide feedback and adjustment for the input content of the live video. In other words, existing solutions do not monitor or adjust live video content, including digital humans. Therefore, improving the quality of live video featuring digital humans is a pressing technical problem that needs to be solved. Summary of the Invention
[0005] This application provides a method and apparatus for adjusting video parameters, which can monitor and adjust the quality monitoring indicators related to digital humans in live video, thereby improving the quality of live video of digital humans.
[0006] In a first aspect, a method for adjusting video parameters is provided, the method comprising: acquiring a first video segment, the first video segment including a digital human, the digital human being being generated by a neural network model; determining video quality monitoring indicators based on the first video segment, the video quality monitoring indicators including at least one of the following: the accuracy of the digital human's lip movements, the synchronization between the digital human's image and sound, and the consistency of the digital human's image; and determining the parameters of the neural network model based on the video quality monitoring indicators, the parameters being used to generate a digital human in a second video segment.
[0007] According to the technical solution provided in this application, by monitoring the monitoring indicators related to digital humans in live video, the model parameters of the generated digital humans can be adjusted based on the video quality reflected by the monitoring indicators, thereby automatically correcting the digital humans in the live video and improving the quality of the live video of digital humans.
[0008] In conjunction with the first aspect, in certain implementations of the first aspect, determining video quality monitoring indicators based on the first video segment includes: determining a first image and a first audio from the first video segment, wherein the first image includes the lip movements of a digital human, and the first audio is the audio of the corresponding frame of the first image in the first video segment; determining a first feature vector based on the first image, wherein the first feature vector is used to describe the lip movement features of the digital human in the first image; determining a second feature vector based on the first audio, wherein the second feature vector is used to describe the audio features of the first audio; and determining the distance between the first feature vector and the second feature vector, wherein the distance is used to represent the accuracy of the lip movements of the digital human.
[0009] According to the above technical solution, by extracting audio features and lip-reading features of digital humans in the same frame and calculating the distance between the audio features and lip-reading features, it is possible to monitor the accuracy of lip-reading in live video.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, the parameters of the neural network model are determined based on video quality monitoring indicators, including: when the distance between the first feature vector and the second feature vector is greater than a first threshold, a first parameter is determined, and the first parameter is used to generate the lip movements of a digital human in the second video segment.
[0011] According to the above technical solution, by automatically adjusting the relevant parameters of the digital human lip-reading generation model through the lip-reading accuracy score reflected by the distance between audio features and lip-reading features, the accuracy of lip-reading in live video can be improved.
[0012] In conjunction with the first aspect, in certain implementations of the first aspect, determining video quality monitoring indicators based on the first video segment includes: determining a time window and a second audio from the first video segment, wherein the time window includes multiple second images corresponding to multiple frames in the first video segment, and the multiple second images include the lip movements of the digital human; determining multiple third feature vectors corresponding to the multiple second images, each third feature vector being used to describe the lip movement features of the digital human in the corresponding second image; determining a fourth feature vector based on the second audio, the fourth feature vector being used to describe the audio features of the second audio; determining a third image from the multiple second images, the third image corresponding to a fifth feature vector, the fifth feature vector being the feature vector with the smallest distance to the fourth feature vector among the multiple third feature vectors; and determining the time difference between the frame corresponding to the third image and the frame corresponding to the second audio, the time difference being used to represent the synchronization degree between the digital human's image and sound.
[0013] According to the above technical solution, by extracting the audio features within a certain frame and the lip-shape features of the digital human in multiple frames within a specific time window, it is possible to find a frame within the time window that has the lip-shape feature with the smallest distance from the audio feature, thereby enabling the monitoring of the audio-visual synchronization of the live video based on the time difference between the frame containing the image and the audio.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, the parameters of the neural network model are determined based on video quality monitoring indicators, including: when the time difference between the frame corresponding to the third image and the frame corresponding to the second audio is greater than a second threshold, the audio of the second video segment is advanced or delayed by the time difference.
[0015] According to the above technical solution, by adjusting the audio of subsequent live videos earlier or later based on the degree of audio-visual synchronization reflected by the time difference, the synchronization between the video and sound of live videos can be improved.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, determining video quality monitoring indicators based on the first video segment includes: determining a fourth image and a fifth image from the first video segment, wherein the fourth image and the fifth image include the facial region of the digital human; determining a sixth feature vector based on the fourth image, wherein the sixth feature vector is used to describe the facial structural features of the digital human in the fourth image; determining a seventh feature vector based on the fifth image, wherein the seventh feature vector is used to describe the facial structural features of the digital human in the fifth image; and determining the distance between the sixth feature vector and the seventh feature vector, wherein the distance is used to represent the consistency of the digital human's image.
[0017] According to the above technical solution, by using the facial structural features of a digital human in a certain frame image as a reference, the similarity between the facial structural features of the digital human in subsequent frames and the facial structural features used as a reference can be calculated, thereby enabling the monitoring of the consistency of the digital human image in the live video based on the above similarity.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, the parameters of the neural network model are determined based on video quality monitoring indicators, including: determining a second parameter when the distance between the sixth feature vector and the seventh feature vector is greater than a third threshold, the second parameter being used to generate the facial structure of a digital human in the second video segment.
[0019] According to the above technical solution, by automatically adjusting the relevant parameters of the digital human generation model based on the consistency of the digital human image reflected by changes in facial features (which can also be expressed as the consistency of the person's identity (id)), the consistency of the digital human image in live video can be improved.
[0020] Secondly, an apparatus for adjusting video parameters is provided, the apparatus comprising: an acquisition module for acquiring a first video segment, the first video segment including a digital human, the digital human being being generated by a neural network model; a monitoring module for determining video quality monitoring indicators based on the first video segment, the video quality monitoring indicators including at least one of the following: the accuracy of the digital human's lip movements, the synchronization between the digital human's image and sound, and the consistency of the digital human's image; and an adjustment module for determining parameters of the neural network model based on the video quality monitoring indicators, the parameters being used to generate a digital human in a second video segment.
[0021] In conjunction with the second aspect, in some implementations of the second aspect, the monitoring module is specifically used to: determine a first image and a first audio from a first video segment, wherein the first image includes the lip movements of a digital human, and the first audio is the audio of the corresponding frame of the first image in the first video segment; determine a first feature vector based on the first image, wherein the first feature vector is used to describe the lip movement features of the digital human in the first image; determine a second feature vector based on the first audio, wherein the second feature vector is used to describe the audio features of the first audio; and determine the distance between the first feature vector and the second feature vector, wherein the distance is used to represent the accuracy of the lip movement of the digital human.
[0022] In conjunction with the second aspect, in some implementations of the second aspect, the adjustment module is specifically used to: determine a first parameter when the distance between the first feature vector and the second feature vector is greater than a first threshold, the first parameter being used to generate the lip movements of a digital human in the second video segment.
[0023] In conjunction with the second aspect, in some implementations of the second aspect, the monitoring module is specifically used to: determine a time window and a second audio from a first video segment, wherein the time window includes multiple second images corresponding to multiple frames in the first video segment, and the multiple second images include the lip movements of the digital human; determine multiple third feature vectors corresponding to the multiple second images, each third feature vector being used to describe the lip movement features of the digital human in the corresponding second image; determine a fourth feature vector based on the second audio, the fourth feature vector being used to describe the audio features of the second audio; determine a third image from the multiple second images, the third image corresponding to a fifth feature vector, the fifth feature vector being the feature vector with the smallest distance to the fourth feature vector among the multiple third feature vectors; and determine the time difference between the frame corresponding to the third image and the frame corresponding to the second audio, the time difference being used to represent the synchronization degree between the digital human's image and sound.
[0024] In conjunction with the second aspect, in some implementations of the second aspect, the adjustment module is specifically used to: advance or delay the audio of the second video segment by a time difference when the time difference between the frame corresponding to the third image and the frame corresponding to the second audio is greater than a second threshold.
[0025] In conjunction with the second aspect, in some implementations of the second aspect, the monitoring module is specifically used to: determine a fourth image and a fifth image from the first video segment, wherein the fourth image and the fifth image include the facial region of the digital human; determine a sixth feature vector based on the fourth image, wherein the sixth feature vector is used to describe the facial structural features of the digital human in the fourth image; determine a seventh feature vector based on the fifth image, wherein the seventh feature vector is used to describe the facial structural features of the digital human in the fifth image; and determine the distance between the sixth feature vector and the seventh feature vector, wherein the distance is used to represent the consistency of the digital human's image.
[0026] In conjunction with the second aspect, in some implementations of the second aspect, the adjustment module is specifically used to: determine the second parameter when the distance between the sixth feature vector and the seventh feature vector is greater than the third threshold. The second parameter is used to generate the facial structure of the digital human in the second video segment.
[0027] Thirdly, a computing device is provided, including a processor and a memory, wherein the memory is used to store instructions, and the processor is used to call and execute the instructions from the memory, causing the computing device to perform the method of the first aspect or any possible implementation thereof.
[0028] Fourthly, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory, wherein the memory is used to store instructions, and the processor is used to call and execute the instructions from the memory, causing the computing device cluster to perform the method in the first aspect or any possible implementation of the first aspect.
[0029] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory, which can be integrated into the processor or exist independently outside the processor.
[0030] Fifthly, a chip is provided that acquires and executes instructions to implement the method in the first aspect or any possible implementation of the first aspect.
[0031] Optionally, as one implementation, the chip includes a processor and a data interface, through which the processor reads instructions stored in the memory and executes the method in the first aspect or any possible implementation of the first aspect.
[0032] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to perform the method in the first aspect or any possible implementation of the first aspect.
[0033] In a sixth aspect, a computer program product containing instructions is provided, which, when executed by a computing device or a cluster of computing devices, causes the computing device or the cluster of computing devices to perform the method described in the first aspect or any possible implementation thereof.
[0034] In a seventh aspect, a computer-readable storage medium is provided, including computer program instructions that, when executed by a computing device or a cluster of computing devices, cause the computing device or the cluster of computing devices to perform the method described in the first aspect or any possible implementation thereof.
[0035] As examples, these computer-readable storage media include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.
[0036] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium. Attached Figure Description
[0037] Figure 1 is a schematic diagram of the system architecture of the live streaming quality monitoring and feedback adjustment system provided in the embodiment of this application.
[0038] Figure 2 is a schematic flowchart of the live video quality monitoring and feedback adjustment method provided in the embodiments of this application.
[0039] Figure 3 is a schematic flowchart of a method for adjusting video parameters provided in an embodiment of this application.
[0040] Figure 4 is a schematic flowchart of a method for monitoring and adjusting the accuracy of digital phrasing provided in an embodiment of this application.
[0041] Figure 5 is a schematic flowchart of a method for monitoring and adjusting the accuracy of digital phrasing provided in an embodiment of this application.
[0042] Figure 6 is a schematic structural block diagram of a device for adjusting video parameters provided in an embodiment of this application.
[0043] Figure 7 is a schematic structural block diagram of a computing device provided in an embodiment of this application.
[0044] Figure 8 is a schematic structural block diagram of a computing device cluster provided in an embodiment of this application.
[0045] Figure 9 is a schematic structural block diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation
[0046] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0047] This application will present various aspects, embodiments, or features relating to systems comprising multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0048] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.
[0049] In the embodiments of this application, "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when the distinction is not emphasized, their intended meanings are consistent.
[0050] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0051] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0052] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0053] Digital humans are digitized figures created using digital technology that possess a human appearance or a resemblance to a human. With the development of artificial intelligence, the intelligence level and the degree of similarity between digital humans and real humans are gradually increasing, enabling digital humans to replace real humans in an growing number of fields. For example, in the live streaming industry, digital humans can replace live streamers, allowing for uninterrupted broadcasts and improving efficiency.
[0054] However, during the live stream, factors such as server computing power, network fluctuations, and weak network conditions may affect the video quality, including but not limited to the clarity of the live video, the degree of audio-visual synchronization, the accuracy of the digital human's lip movements, and the consistency of the image, which may affect the live stream effect.
[0055] Traditional live video quality monitoring and adjustment technologies primarily focus on the architectural design of the live streaming platform and / or video streaming, but neglect content quality monitoring. For example, monitoring common video data-related information such as bitrate and resolution of the live video stream can generate monitoring results and provide them to operations and maintenance personnel, thereby maintaining the bitrate stability and image clarity of the live video. However, these technologies cannot provide feedback and adjustment to the input content of the live video. In other words, the monitoring metrics of traditional solutions do not involve the high-level semantic information contained in the video content, and therefore cannot monitor or adjust for anomalies in digital figures within the live content.
[0056] Therefore, improving the quality of live-streamed videos featuring digital humans has become an urgent technical problem to be solved.
[0057] This application provides a method and apparatus for adjusting video parameters, which can monitor and adjust the quality monitoring indicators related to digital humans in live video, thereby improving the quality of live video of digital humans.
[0058] Figure 1 is a schematic diagram of the system architecture of the live streaming quality monitoring and feedback adjustment system provided in an embodiment of this application. As shown in Figure 1, live streaming settings parameters are input into the digital human live streaming system to generate a live video feed containing a digital human. These live streaming settings parameters may include conventional video data-related parameters, such as resolution settings, frame rate settings, and bitrate settings, and may also include parameters related to the generation of the digital human, such as model parameters for generating facial features, model parameters for generating lip shapes, and audio-visual synchronization settings for the digital human's image and sound. Optionally, model parameters may include training parameters and / or inference parameters of the model.
[0059] The generated live video feed is input into a quality monitoring system to obtain quality monitoring indicators for the live video. These indicators can include conventional video data-related metrics, such as image clarity and bitrate stability, as well as metrics related to the digital humans in the live content, such as the accuracy of the digital humans' lip movements, the synchronization between the digital humans' image and sound, and the consistency of the digital humans' appearance. The consistency of the digital humans' appearance can be represented by the consistency of their identity IDs.
[0060] Based on the aforementioned quality monitoring indicators, the system can determine whether adjustments to the input live streaming settings parameters are necessary. For example, if all quality monitoring indicators are normal, no adjustments are made, and the live streaming video is output according to the current live streaming settings parameters. If one or more quality monitoring indicators are abnormal, the live streaming settings parameters can be automatically adjusted through the feedback adjustment system. The new live streaming settings parameters are then input into the digital human live streaming system, and subsequent output live streaming videos are generated according to the new live streaming settings parameters, thereby promptly correcting the parameters and / or content of the live streaming video and maintaining stable live streaming video quality.
[0061] Figure 2 is a schematic flowchart of the live video quality monitoring and feedback adjustment method provided in this application embodiment. As shown in Figure 2, the live video stream may include a video stream and an audio stream. The video stream may include multiple images, each image corresponding to a frame in the live video. The video stream and audio stream can respectively extract image features and audio features through corresponding neural network models. The extracted video features and audio features can be jointly input into a third neural network model to obtain video quality monitoring indicators related to the digital human. For example, video features and audio features can be jointly input into a lip-sync accuracy evaluation model to obtain a lip-sync accuracy score, thereby adjusting the parameters of the digital human's lip-sync generation model based on the feedback from the lip-sync accuracy score. As another example, video features and audio features can be jointly input into another lip-sync accuracy evaluation model to obtain an audio-visual synchronization score, thereby adjusting the digital human's audio-visual alignment parameters based on the feedback from the audio-visual synchronization score. Furthermore, video features and / or audio features can also be used individually to determine video quality monitoring indicators related to the digital human. For example, video features can be used individually to obtain a person ID consistency score, thereby adjusting the parameters of the digital human's facial feature generation model based on the person ID consistency score feedback.
[0062] The method for adjusting video parameters according to this application will be described in detail below with reference to Figure 3. Optionally, this method can be applied to the system shown in Figure 1. As shown in Figure 3, the method includes the following steps.
[0063] Step S310: Obtain the first video segment.
[0064] Specifically, the first video segment can be a live video segment containing a digital human, wherein the digital human can be generated by a neural network model. For example, in step S310, the quality monitoring system can acquire the live video segment currently being streamed, or the quality monitoring system can acquire the live video segment that is about to be streamed, to obtain quality monitoring parameters related to the digital human in the video based on the aforementioned live video segment.
[0065] Step S320: Determine the video quality monitoring indicators based on the first video segment.
[0066] For example, in step S320, the quality monitoring system can obtain quality monitoring parameters related to the digital human in the video based on the live video segment obtained in S310. Video quality monitoring indicators may include, but are not limited to, at least one of the following: the accuracy of the digital human's lip movements in the video segment (hereinafter referred to as lip-sync accuracy), the synchronization between the digital human's image and sound in the video segment (hereinafter referred to as audio-visual synchronization), and the consistency of the digital human's image in the video segment. Specifically, lip-sync accuracy refers to the degree of matching between the lip movements of the digital human in a frame of the video stream and the audio in the same frame of the audio stream; audio-visual synchronization refers to the time difference between the audio in a frame of the audio stream and the most matching frame of the video stream; and image consistency refers to the similarity between the image of the digital human, including facial structure, in a frame of the video stream and the image of the digital human in a reference frame image. Optionally, image consistency can be represented by the consistency of the human's ID.
[0067] Optionally, the aforementioned video quality monitoring indicators can be obtained through a neural network model based on image features extracted from the video stream and / or audio features extracted from the audio stream. It should be understood that the feature types and / or neural network models used to obtain different video quality monitoring indicators may differ. The following sections will illustrate the acquisition of several video quality monitoring indicators related to digital humans with specific examples, and will not elaborate further here.
[0068] Step S330: Determine the parameters of the neural network model based on the video quality monitoring indicators.
[0069] For example, in step S330, the quality monitoring system can determine the setting parameters of the second video segment based on the quality monitoring indicators of the first video segment determined in S320. Optionally, the second video segment can be a live video segment that is time-sequentially located after the first video segment. In other words, the quality monitoring system can adjust the setting parameters of the next video segment to be live-streamed based on the quality of the currently being live-streamed video.
[0070] The adjustment of parameters can include, but is not limited to, adjustments to the parameters of the neural network model used to generate the digital human in the second video segment. For example, the training parameters of the neural network model used to generate the digital human can be adjusted, including but not limited to deleting dirty training data, modifying the batch size, and modifying the weights of the loss function; the inference parameters of the neural network model used to generate the digital human can also be adjusted, including but not limited to frame selection and time window selection. Furthermore, the parameters can also include parameters related to live video streaming, such as resolution settings and frame rate settings; this application does not specifically limit this.
[0071] In some possible implementations, video quality monitoring indicators can be quantified into scores, such as lip-sync accuracy scores, audio-visual synchronization scores, and character ID consistency scores. The settings parameters can then be adjusted based on the relationship between these scores and specific thresholds. For example, a higher score indicates higher quality in that aspect of the live video. When a score is greater than or equal to the corresponding threshold, the quality monitoring system can maintain the current settings parameters; that is, the same parameter settings can be used to generate the second video segment as the first video segment. When a score is less than the specific threshold, the feedback adjustment system can automatically adjust the relevant settings parameters; that is, the settings parameters used to generate the second video segment may not be exactly the same as those used for the first video segment.
[0072] The technical solution of this application embodiment monitors the monitoring indicators related to digital humans in live videos, and can adjust the model parameters of the generated digital humans based on the video quality reflected by the monitoring indicators, thereby automatically correcting the digital humans in the live videos and improving the quality of the digital human live videos.
[0073] The following describes, with reference to three specific embodiments, the methods for monitoring and adjusting the lip-sync accuracy, audio-visual synchronization, and character ID consistency of digital humans in live video provided by the technical solution of this application.
[0074] In some possible implementations, in step S320 above, the accuracy of the lip movements of the digital human in the live video can be monitored by extracting audio features within the same frame and lip-sync features in the image, and calculating the distance between the audio features and lip-sync features. Figure 4 shows a schematic flowchart of the method for monitoring and adjusting the accuracy of digital lip movements provided in this application embodiment. As shown in Figure 4, for the input video segment, the video part and the audio part are extracted separately. Both the video part and the audio part can be divided into the smallest unit in the time domain according to the pre-trained frame rate. For example, the neural network model used for feature extraction can be trained at a frame rate of 25, so 25 images and corresponding 25 audio segments can be extracted from each second of the input video segment. For the video part, each frame can be regarded as an image, so the lip-sync feature vector can be extracted from each image by the neural network model. The lip-sync feature vector is used to describe the lip-sync features of the digital human in the image. For the audio part, the audio within each frame can be regarded as an independent audio segment, so the audio feature vector can be extracted from each audio segment by the neural network model. Optionally, for each audio segment, Hubert audio features can be extracted first, followed by feature vector extraction from the Hubert audio features, thereby reducing the dimensionality of the final extracted audio feature vector. After extracting the lip-sync feature vector and audio feature vector respectively, the lip-sync feature vector of the image and the audio feature vector of the audio within the same frame can be fed into a third neural network model for difference calculation to obtain the distance between the lip-sync feature vector and the audio feature vector. This distance can be used to represent the lip-sync accuracy of the digital human within that frame; the closer the distance between the lip-sync feature vector and the audio feature vector, the higher the accuracy of the digital human's lip-sync; the farther the distance between the lip-sync feature vector and the audio feature vector, the lower the accuracy of the digital human's lip-sync.
[0075] Optionally, the neural network model for extracting and comparing feature vectors described above can be obtained by training a large number of audio-visual synchronized broadcast videos as training data.
[0076] Optionally, after calculating the distance between the lip-shape feature vector and the audio feature vector, this distance can be quantified as a lip-shape accuracy score. If the lip-shape accuracy score is less than a certain threshold, or the distance between the lip-shape feature vector and the audio feature vector is greater than a first threshold, the parameters used to generate the digital human's lip movements can be automatically adjusted. These adjustments can include retraining and / or fine-tuning the lip-shape generation model. For example, the training parameters of the lip-shape generation model can be adjusted for retraining, including but not limited to automatic audio-visual alignment of training data, deletion of dirty training data, batch size modification, and modification of loss function weights; the inference parameters of the lip-shape generation model can also be adjusted for fine-tuning, including but not limited to setting the length of the sound segment for each inference and setting the skip-size for inaccurate audio. Through these methods, the lip-shape accuracy of digital humans in live videos can be improved.
[0077] In other possible implementations, in step S320 above, audio features within a certain frame and lip-sync features of the digital human in multiple frames within a specific time window can be extracted. A frame within that time window with the lip-sync feature closest to the audio feature can then be found, thereby monitoring the audio-visual synchronization of the live video based on the time difference between the frames containing the images and the audio. Figure 5 shows a schematic flowchart of the method for monitoring and adjusting the accuracy of digital lip-sync provided in this application embodiment. As shown in Figure 5, the input video segment is first processed into a video part and an audio part, and lip-sync feature vectors in the image and audio feature vectors in the audio are extracted frame by frame. Optionally, the above feature extraction process can be the same as the feature extraction steps in the embodiment shown in Figure 4. For specific implementation details, please refer to the previous description of Figure 4, which will not be repeated here. After extracting the feature vectors, the audio feature vector of a certain frame can be used as a reference to find a frame within a specific time window such that the distance between the lip-sync feature vector of that image and the reference audio feature vector is minimized within that time window. Then, the specific time difference between the frame containing the reference audio and the frame containing the found image can be quantitatively calculated to represent the degree of audio-visual asynchrony in the live video.
[0078] For example, a 21-frame window can be selected from a video clip, and the audio feature vector of the middle frame (frame 0) can be used as a reference. Lip-sync feature vectors can be extracted from frame 0 and the 10 frames before and after it, and the distance to the reference audio feature vector can be calculated for each. This allows us to determine the lip-sync feature vector with the smallest distance from the audio vector among the 21 lip-sync feature vectors, and the corresponding frame. Assuming the lip-sync feature vector of frame 5 in the window has the smallest distance from the reference audio feature vector, the audio-visual asynchrony is 5 frames. Depending on the frame rate used, this 5-frame difference can be converted into a time difference in seconds.
[0079] Optionally, after calculating the time difference between the frames corresponding to the lip-sync feature vector and the frames corresponding to the audio feature vector, this time difference can be quantified as an audio-visual synchronization score. If the audio-visual synchronization score is less than a certain threshold, or if the time difference between the frames corresponding to the lip-sync feature vector and the frames corresponding to the audio feature vector is greater than a second threshold, the parameters used to generate the digital lip-sync can be automatically adjusted. For example, if the found lip-sync feature vector is 5 frames ahead of the reference audio feature vector, the audio track of the subsequently generated live video will be advanced by the time corresponding to 5 frames; if the found lip-sync feature vector is 4 frames behind the reference audio feature vector, the audio track of the subsequently generated live video will be delayed by the time corresponding to 4 frames. This approach improves the synchronization between the video and audio in live videos.
[0080] In other possible implementations, in step S330 above, the facial structural features of the digital human in a certain frame image can be used as a reference to calculate the similarity between the facial structural features of the digital human in subsequent frames and the reference facial structural features, thereby monitoring the consistency of the digital human's image in the live video based on the aforementioned similarity. Specifically, one frame image from the input video segment can be selected as a reference, and the facial structural feature vector in that frame image can be extracted. This facial structural feature vector can be used to describe the facial structural features of the digital human in the image. Then, facial structural feature vectors can be extracted from other frames of the input video segment and compared with the facial structural feature vector of the reference image to calculate the distance between them. This distance can be used to represent the consistency of the digital human's image between two frames; the closer the distance between the two facial structural feature vectors, the higher the consistency of the digital human's image between the two frames; the farther the distance between the two facial structural feature vectors, the lower the consistency of the digital human's image between the two frames.
[0081] Optionally, after calculating the distance between two facial structure feature vectors, this distance can be quantified as a person ID consistency score. If the person ID consistency score is less than a certain threshold, or the distance between the two facial structure feature vectors is greater than a third threshold, the parameters used to generate the digital human facial structure can be automatically adjusted. These adjustments can include retraining and / or fine-tuning the facial structure generation model. For example, the training parameters of the facial structure generation model can be adjusted for retraining, including but not limited to beard region parameter settings, face mask region adjustments, neck mask region adjustments, and loss function weight modifications; the inference parameters of the facial structure generation model can also be adjusted for fine-tuning, including but not limited to reference frame selection settings and face mask region adjustments. Through these methods, the consistency of the digital human image in live video can be improved.
[0082] The above description, with reference to Figures 1 to 5, illustrates an embodiment of the method for adjusting video parameters provided in this application. The following description, with reference to Figures 6 to 9, illustrates an embodiment of the apparatus for adjusting video parameters provided in this application.
[0083] Figure 6 shows a schematic structural diagram of a device 600 for adjusting video parameters provided in an embodiment of this application.
[0084] As shown in Figure 6, the device 600 for adjusting video parameters includes: an acquisition module 610, a monitoring module 620, and an adjustment module 630.
[0085] Specifically, the acquisition module 610 is used to acquire a first video segment, which includes a digital human generated by a neural network model.
[0086] Specifically, the monitoring module 620 is used to determine video quality monitoring indicators based on the first video segment. The video quality monitoring indicators include at least one of the following: the accuracy of the digital human's lip movements, the synchronization between the digital human's image and sound, and the consistency of the digital human's appearance.
[0087] Optionally, the monitoring module 620 is specifically configured to determine a first image and a first audio from a first video segment, wherein the first image includes the lip movements of a digital human, and the first audio is the audio of the corresponding frame of the first image in the first video segment; determine a first feature vector based on the first image, wherein the first feature vector is used to describe the lip movement features of the digital human in the first image; determine a second feature vector based on the first audio, wherein the second feature vector is used to describe the audio features of the first audio; and determine the distance between the first feature vector and the second feature vector, wherein the distance is used to represent the accuracy of the lip movement of the digital human.
[0088] Optionally, the monitoring module 620 is specifically used to determine a time window and a second audio from a first video segment. The time window includes multiple second images corresponding to multiple frames in the first video segment, and the multiple second images include the lip movements of the digital human. Based on the multiple second images, multiple third feature vectors corresponding to the multiple second images are determined, and each third feature vector is used to describe the lip movement features of the digital human in the corresponding second image. Based on the second audio, a fourth feature vector is determined, and the fourth feature vector is used to describe the audio features of the second audio. A third image is determined from the multiple second images, and the third image corresponds to a fifth feature vector, which is the feature vector with the smallest distance to the fourth feature vector among the multiple third feature vectors. The time difference between the frame corresponding to the third image and the frame corresponding to the second audio is determined, and the time difference is used to represent the synchronization degree between the digital human's image and sound.
[0089] Optionally, the monitoring module 620 is specifically used to determine a fourth image and a fifth image from the first video segment, the fourth image and the fifth image including the facial region of the digital human; determine a sixth feature vector based on the fourth image, the sixth feature vector being used to describe the facial structural features of the digital human in the fourth image; determine a seventh feature vector based on the fifth image, the seventh feature vector being used to describe the facial structural features of the digital human in the fifth image; and determine the distance between the sixth feature vector and the seventh feature vector, the distance being used to represent the consistency of the digital human's image.
[0090] Specifically, the adjustment module 630 is used to determine the parameters of the neural network model based on video quality monitoring indicators. These parameters are used to generate a digital human in the second video segment.
[0091] Optionally, the adjustment module 630 is specifically used to determine a first parameter when the distance between the first feature vector and the second feature vector is greater than a first threshold. The first parameter is used to generate the lip movements of a digital human in the second video segment.
[0092] Optionally, the adjustment module 630 is specifically used to advance or delay the audio of the second video segment by a time difference when the time difference between the frame corresponding to the third image and the frame corresponding to the second audio is greater than a second threshold.
[0093] Optionally, the adjustment module 630 is specifically used to determine a second parameter when the distance between the sixth feature vector and the seventh feature vector is greater than a third threshold. The second parameter is used to generate the facial structure of the digital human in the second video segment.
[0094] All of the above modules can be implemented in software or hardware. For example, the implementation of monitoring module 620 will be described below. Similarly, the implementation of acquisition module 610 and adjustment module 630 can be referenced from the implementation of monitoring module 620.
[0095] As an example of a software functional unit, the monitoring module 620 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the monitoring module 620 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0096] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0097] As an example of a hardware functional unit, the monitoring module 620 may include at least one computing device, such as a server. Alternatively, the matching module 520 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0098] The monitoring module 620 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the monitoring module 620 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the monitoring module 620 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0099] It should be noted that, in other embodiments, the acquisition module 610, the monitoring module 620, and the adjustment module 630 can be used to execute any step in the above-described method for adjusting video parameters. The steps implemented by the acquisition module 610, the monitoring module 620, and the adjustment module 630 can be specified as needed. By implementing different steps in the above-described method for adjusting video parameters through the acquisition module 610, the monitoring module 620, and the adjustment module 630, all functions of the device 600 can be realized.
[0100] This application also provides a computing device 100. As shown in FIG7, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.
[0101] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 7, but this does not imply that there is only one bus or one type of bus. Bus 102 can include pathways for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108).
[0102] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0103] The memory 106 may include volatile memory, such as random access memory (RAM). The memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0104] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned acquisition module, monitoring module, and adjustment module, thereby realizing the method for adjusting video parameters. In other words, the memory 106 stores instructions for executing the method for adjusting video parameters.
[0105] The communication interface 108 uses a command distribution module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 100 and other devices or communication networks.
[0106] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0107] As shown in Figure 8, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for performing the above-described method of adjusting video parameters.
[0108] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the above-described method of adjusting video parameters. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the above-described method of adjusting video parameters.
[0109] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned video parameter adjustment device. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules among the acquisition module, monitoring module, and adjustment module.
[0110] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 illustrates one possible implementation. As shown in Figure 9, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A stores instructions for performing the functions of the acquisition module and the monitoring module. Simultaneously, the memory 106 in computing device 100B stores instructions for performing the functions of the adjustment module.
[0111] It should be understood that the functions of computing device 100A shown in Figure 9 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.
[0112] This application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in the memory through the data interface to execute the above-described method for adjusting video parameters.
[0113] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the above-described method for adjusting video parameters.
[0114] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform the above-described method for adjusting video parameters.
[0115] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0116] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A method for adjusting video parameters, characterized in that, include: Acquire a first video clip, the first video clip including a digital human, the digital human being being generated by a neural network model; Based on the first video segment, video quality monitoring indicators are determined, including at least one of the following: the accuracy of the lip movements of the digital human, the synchronization between the image and sound of the digital human, and the consistency of the image of the digital human. Based on the video quality monitoring indicators, the parameters of the neural network model are determined, and these parameters are used to generate the digital human in the second video segment.
2. The method according to claim 1, characterized in that, The step of determining video quality monitoring indicators based on the first video segment includes: A first image and a first audio are determined from the first video segment, wherein the first image includes the lip movements of the digital human, and the first audio is the audio of the corresponding frame of the first image in the first video segment; Based on the first image, a first feature vector is determined, which is used to describe the lip-shape features of the digital human in the first image; Based on the first audio, a second feature vector is determined, which is used to describe the audio features of the first audio. The distance between the first feature vector and the second feature vector is determined, and the distance is used to represent the lip-reading accuracy of the digital human.
3. The method according to claim 2, characterized in that, The step of determining the parameters of the neural network model based on the video quality monitoring indicators includes: If the distance between the first feature vector and the second feature vector is greater than a first threshold, a first parameter is determined, which is used to generate the lip movements of the digital human in the second video segment.
4. The method according to any one of claims 1 to 3, characterized in that, The step of determining video quality monitoring indicators based on the first video segment includes: A time window and a second audio are determined from the first video segment. The time window includes multiple second images corresponding to multiple frames in the first video segment, and the multiple second images include the lip movements of the digital human. Based on the plurality of second images, a plurality of third feature vectors corresponding to the plurality of second images are determined, and each third feature vector is used to describe the lip shape features of the digital human in the corresponding second image; Based on the second audio, a fourth feature vector is determined, which is used to describe the audio features of the second audio. A third image is determined from the plurality of second images, the third image corresponding to a fifth feature vector, the fifth feature vector being the feature vector with the smallest distance to the fourth feature vector among the plurality of third feature vectors; The time difference between the frame corresponding to the third image and the frame corresponding to the second audio is determined, and the time difference is used to represent the synchronization degree between the image and sound of the digital human.
5. The method according to claim 4, characterized in that, The step of determining the parameters of the neural network model based on the video quality monitoring indicators includes: If the time difference between the frame corresponding to the third image and the frame corresponding to the second audio is greater than a second threshold, the audio of the second video segment will be advanced or delayed by the time difference.
6. The method according to any one of claims 1 to 5, characterized in that, The step of determining video quality monitoring indicators based on the first video segment includes: A fourth image and a fifth image are determined from the first video segment, the fourth image and the fifth image including the facial region of the digital human; Based on the fourth image, a sixth feature vector is determined, which is used to describe the facial structure features of the digital human in the fourth image; Based on the fifth image, a seventh feature vector is determined, which is used to describe the facial structure features of the digital human in the fifth image; The distance between the sixth feature vector and the seventh feature vector is determined, and the distance is used to represent the consistency of the digital human's image.
7. The method according to claim 6, characterized in that, The step of determining the parameters of the neural network model based on the video quality monitoring indicators includes: If the distance between the sixth feature vector and the seventh feature vector is greater than a third threshold, a second parameter is determined, which is used to generate the facial structure of the digital human in the second video segment.
8. A device for adjusting video parameters, characterized in that, include: An acquisition module is used to acquire a first video segment, the first video segment including a digital human, the digital human being being generated by a neural network model; The monitoring module is used to determine video quality monitoring indicators based on the first video segment. The video quality monitoring indicators include at least one of the following: the accuracy of the lip movements of the digital human, the synchronization between the image and sound of the digital human, and the consistency of the image of the digital human. An adjustment module is used to determine the parameters of the neural network model based on the video quality monitoring indicators, and the parameters are used to generate the digital human in the second video segment.
9. The apparatus according to claim 8, characterized in that, The monitoring module is used for: A first image and a first audio are determined from the first video segment, wherein the first image includes the lip movements of the digital human, and the first audio is the audio of the corresponding frame of the first image in the first video segment; Based on the first image, a first feature vector is determined, which is used to describe the lip-shape features of the digital human in the first image; Based on the first audio, a second feature vector is determined, which is used to describe the audio features of the first audio. The distance between the first feature vector and the second feature vector is determined, and the distance is used to represent the lip-reading accuracy of the digital human.
10. The apparatus according to claim 9, characterized in that, The adjustment module is used for: If the distance between the first feature vector and the second feature vector is greater than a first threshold, a first parameter is determined, which is used to generate the lip movements of the digital human in the second video segment.
11. The apparatus according to any one of claims 8 to 10, characterized in that, The monitoring module is used for: A time window and a second audio are determined from the first video segment. The time window includes multiple second images corresponding to multiple frames in the first video segment, and the multiple second images include the lip movements of the digital human. Based on the plurality of second images, a plurality of third feature vectors corresponding to the plurality of second images are determined, and each third feature vector is used to describe the lip shape features of the digital human in the corresponding second image; Based on the second audio, a fourth feature vector is determined, which is used to describe the audio features of the second audio. A third image is determined from the plurality of second images, the third image corresponding to a fifth feature vector, the fifth feature vector being the feature vector with the smallest distance to the fourth feature vector among the plurality of third feature vectors; The time difference between the frame corresponding to the third image and the frame corresponding to the second audio is determined, and the time difference is used to represent the synchronization degree between the image and sound of the digital human.
12. The apparatus according to claim 11, characterized in that, The adjustment module is used for: If the time difference between the frame corresponding to the third image and the frame corresponding to the second audio is greater than a second threshold, the audio of the second video segment will be advanced or delayed by the time difference.
13. The apparatus according to any one of claims 8 to 12, characterized in that, The monitoring module is used for: A fourth image and a fifth image are determined from the first video segment, the fourth image and the fifth image including the facial region of the digital human; Based on the fourth image, a sixth feature vector is determined, which is used to describe the facial structure features of the digital human in the fourth image; Based on the fifth image, a seventh feature vector is determined, which is used to describe the facial structure features of the digital human in the fifth image; The distance between the sixth feature vector and the seventh feature vector is determined, and the distance is used to represent the consistency of the digital human's image.
14. The apparatus according to claim 13, characterized in that, The adjustment module is used for: If the distance between the sixth feature vector and the seventh feature vector is greater than a third threshold, a second parameter is determined, which is used to generate the facial structure of the digital human in the second video segment.
15. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 7.
16. A computer program product, characterized in that, Includes instructions that, when executed by a computing device or a cluster of computing devices, cause the computing device or the cluster of computing devices to perform the method as described in any one of claims 1 to 7.
17. A computer-readable storage medium, characterized in that, Includes computer program instructions that, when executed by a computing device or cluster of computing devices, cause the computing device or cluster of computing devices to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Digital human video generation method and device, electronic equipment and storage medium
CN113987269A
Audio and video synchronization method, system and device of AI digital human in live broadcast and medium
CN115720275A
Training evaluation model and method and device for evaluating video quality
CN117012228A
Automatic alignment method and system for digital population types
CN118612490A
Computing device and method for realistic visualization of digital human
US20240193824A1
Cited By
A method, apparatus, storage medium and electronic device for driving a digital human
CN122176130A