Video platform video processing method and video platform
By adjusting the timing of audio frames and video frames to generate wrong frequency videos, the problem of virtual human generation technology being maliciously used is solved, and the effect of distinguishing virtual character videos and improving user experience is achieved.
Patent Information
- Application Number
- CN202510111535.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
Existing virtual life generation technology may be maliciously used, resulting in ethical and security issues in online audio and video interactions, such as online fraud and virtual identity impersonation.
By obtaining pre-stored error frequency values and original videos, adjusting the timing of audio frames and video frames, generating error frequency videos, and fixing error frequency videos when users need them to restore the original video.
By generating wrong frequency videos, the mismatch between audio frames and video frames is increased, resulting in video discomfort, thereby distinguishing virtual character videos and avoiding malicious use; at the same time, repairing the video after the user is informed, improving the viewing experience.
Smart Images

Figure CN119946385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a video platform video processing method and a video platform. Background Art
[0002] With the continuous advancement of virtual human generation technology, digital humans and simulated humans have been widely used in entertainment, education, medical care and other fields.
[0003] The publicly available virtual human generation technologies, such as StyleGAN face generation technology, DeepFake video generation technology, and Wav2Lip audio and video synchronization generation technology, mostly focus on the optimization of voice-driven facial expression generation algorithms, aiming to improve naturalness and synchronization, thereby improving the simulation of virtual characters.
[0004] However, this means that the current virtual human generation technology may be used maliciously, which may easily lead to ethical and security issues when conducting online audio and video interactions through virtual simulated characters, such as online fraud and virtual identity impersonation.
[0005] Therefore, it is necessary to provide a new video platform video processing method and video platform to solve the above problems existing in the prior art. Summary of the invention
[0006] The purpose of the present invention is to provide a video platform video processing method and a video platform, which can distinguish videos as virtual character videos, thereby preventing virtual human generation technology from being maliciously used.
[0007] To achieve the above object, the video platform video processing method of the present invention comprises the following steps:
[0008] S1: Obtain pre-stored error frequency value and original video;
[0009] S2: adjusting the timing of at least one of the audio frame and the video frame in the original video according to the error frequency value to generate an error frequency video and provide it to the user;
[0010] S3: When the user needs to restore the video, the wrong video is repaired to obtain the original video and provide it to the user.
[0011] The beneficial effects of the video platform video processing method are: obtaining an error frequency value and an original video, adjusting the timing of at least one of the audio frame and the video frame in the original video according to the error frequency value to generate an error frequency video and provide it to the user, so that the audio frame in the error frequency video is not synchronized with the corresponding video frame, that is, there is an error frequency between the video frame and the audio frame, so that the error frequency video has an uncomfortable feeling and can be distinguished as a virtual character video, thereby avoiding the malicious use of virtual human generation technology, and when the user needs to restore the video, the error frequency video is repaired to obtain the original video and provide it to the user. After the user knows that it is a virtual character video, it is repaired to the original video for the user to watch, thereby improving the viewing experience.
[0012] Optionally, step S2 further includes encrypting the frequency error value using a preset public key to obtain an encrypted frequency error value.
[0013] Optionally, step S3 further includes decrypting the encrypted error frequency value by using a private key corresponding to the public key to obtain the error frequency value.
[0014] Optionally, in step S3, the frequency-error video is repaired according to the frequency-error value.
[0015] Optionally, in step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the original video is adjusted according to the frequency error value to generate a frequency error video and provide it to the user.
[0016] Optionally, the video platform video processing method also includes an audio recognition step, which includes identifying guiding words and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of the guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address.
[0017] Optionally, in step S3, after repairing the wrong frequency video, it also includes adding annotation information to the original video obtained by repairing the wrong frequency video to indicate that the original video obtained by repairing the wrong frequency video is a virtual character video.
[0018] Optionally, in step S1, the error frequency value is obtained from encryption hardware.
[0019] The present invention also provides a video platform, including an acquisition unit, a frequency error adjustment unit and a frequency error adjustment recovery unit, wherein the acquisition unit is used to acquire a frequency error value and an original video, the frequency error adjustment unit is used to adjust the timing of at least one of the audio frames and video frames in the original video according to the frequency error value to generate a frequency error video and provide it to a user, and the frequency error adjustment recovery unit is used to repair the frequency error video to obtain the original video and provide it to the user.
[0020] The beneficial effect of the video platform is that: the acquisition unit acquires the error frequency value and the original video, and the error frequency adjustment unit adjusts the timing of at least one of the audio frame and the video frame in the original video according to the error frequency value to generate an error frequency video and provide it to the user, so that the audio frame in the error frequency video is not synchronized with the corresponding video frame, that is, there is an error frequency between the video frame and the audio frame, so that the error frequency video has an uncomfortable feeling and can be distinguished as a virtual character video, thereby avoiding the malicious use of virtual human generation technology, and when the user needs to restore the video, the error frequency adjustment and recovery unit repairs the error frequency video to obtain the original video and provide it to the user. After the user knows that it is a virtual character video, it is restored to the original video for the user to watch, thereby improving the viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flowchart of a video platform video processing method in some embodiments of the present invention;
[0022] Figure 2 A schematic diagram of the timing of coordination of audio frames and video frames in an original video in some embodiments of the present invention;
[0023] Figure 3 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;
[0024] Figure 4 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;
[0025] Figure 5 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;
[0026] Figure 6 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;
[0027] Figure 7 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;
[0028] Figure 8 In some embodiments of the present invention, Figure 2 Schematic diagram of the coordination timing of audio frames and video frames in the new video generated from the original video. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0030] In view of the problems existing in the prior art, an embodiment of the present invention provides a video processing method for a video platform. Figure 1 , the video platform video processing method comprises the following steps:
[0031] S1: Obtain pre-stored error frequency value and original video;
[0032] S2: adjusting the timing of at least one of the audio frame and the video frame in the original video according to the error frequency value to generate an error frequency video and provide it to the user;
[0033] S3: When the user needs to restore the video, the wrong video is repaired to obtain the original video and provide it to the user.
[0034] In some embodiments, when executing step S1, the same or different error frequency values can be continuously obtained, so that when executing step S2, the same or different timing adjustments can be made to different time periods of the audio frames and video frames in the original video.
[0035] Figure 2 FIG. 1 is a schematic diagram of the timing of the coordination of audio frames and video frames in the original video in some embodiments of the present invention. Figure 2 , the original video includes video frame A, video frame B, video frame C, video frame D, video frame E, video frame F, video frame G, video frame H, video frame I, video frame J, video frame K, video frame L, video frame M, video frame N, video frame O, video frame P, video frame Q, video frame R, audio frame a, audio frame b, audio frame c, audio frame d, audio frame e, audio frame f, audio frame g, audio frame h, audio frame i, audio frame j, audio frame k, audio frame l, audio frame m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r.
[0036] Reference Figure 2, video frame A corresponds to audio frame a, video frame B corresponds to audio frame b, video frame C corresponds to audio frame c, video frame D corresponds to audio frame d, video frame E corresponds to audio frame e, video frame F corresponds to audio frame f, video frame G corresponds to audio frame g, video frame H corresponds to audio frame h, video frame I corresponds to audio frame i, video frame J corresponds to audio frame j, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.
[0037] Figure 3 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 3 , taking the generated error frequency value of -2 as an example, the audio frame is delayed as a whole, and video frame A and video frame B have no corresponding audio frames, that is, no sound is emitted when video frame A and video frame B are played; video frame C corresponds to audio frame a, video frame D corresponds to audio frame b, video frame E corresponds to audio frame c, video frame F corresponds to audio frame d, video frame G corresponds to audio frame e, video frame H corresponds to audio frame f, video frame I corresponds to audio frame g, video frame J corresponds to audio frame h, video frame K corresponds to audio frame i, video frame L corresponds to audio frame j, video frame M corresponds to audio frame k, video frame N corresponds to audio frame l, video frame O corresponds to audio frame m, video frame P corresponds to audio frame n, video frame Q corresponds to audio frame o, and video frame R corresponds to audio frame p.
[0038] Reference Figure 3 Due to the overall delay of the audio frames, the audio frames q and r have no corresponding video frames. The audio frames q and r can be played without video images, or the audio frames q and r can be deleted.
[0039] Figure 4 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 4, taking the generated error frequency value of -6 as an example, the audio frame is delayed as a whole, and video frame A, video frame B, video frame C, video frame D, video frame E and video frame F have no corresponding audio frames, that is, no sound is emitted when playing video frame A, video frame B, video frame C, video frame D, video frame E and video frame F; video frame G corresponds to audio frame a, video frame H corresponds to audio frame b, video frame I corresponds to audio frame c, video frame J corresponds to audio frame d, video frame K corresponds to audio frame e, video frame L corresponds to audio frame f, video frame M corresponds to audio frame g, video frame N corresponds to audio frame h, video frame O corresponds to audio frame i, video frame P corresponds to audio frame j, video frame Q corresponds to audio frame k, and video frame R corresponds to audio frame l.
[0040] Reference Figure 4 Due to the overall delay of audio frames, audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r have no corresponding video frames. Audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r can be played without video images, and audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r can also be deleted.
[0041] Figure 5 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 5 , taking the generated error frequency value of 2 as an example, the video frame is delayed as a whole, and audio frame a and audio frame b have no corresponding video frames, that is, no video image is played when audio frame a and audio frame b are played; video frame A corresponds to audio frame c, video frame B corresponds to audio frame d, video frame C corresponds to audio frame e, video frame D corresponds to audio frame f, video frame E corresponds to audio frame g, video frame F corresponds to audio frame h, video frame G corresponds to audio frame i, video frame H corresponds to audio frame j, video frame I corresponds to audio frame k, video frame J corresponds to audio frame l, video frame K corresponds to audio frame m, video frame L corresponds to audio frame n, video frame M corresponds to audio frame o, video frame N corresponds to audio frame p, video frame O corresponds to audio frame q, and video frame P corresponds to audio frame r.
[0042] Reference Figure 5 Due to the overall delay of the video frames, video frames Q and R have no corresponding audio frames, and no sound is emitted when playing video frames Q and R.
[0043] Figure 6 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 6, taking the generated error frequency value of 6 as an example, the video frame is delayed as a whole, and audio frame a, audio frame b, audio frame c, audio frame d, audio frame e and audio frame f have no corresponding video frames, that is, no video image is played when audio frame a, audio frame b, audio frame c, audio frame d, audio frame e and audio frame f are played; video frame A corresponds to audio frame g, video frame B corresponds to audio frame h, video frame C corresponds to audio frame i, video frame D corresponds to audio frame j, video frame E corresponds to audio frame k, video frame F corresponds to audio frame l, video frame G corresponds to audio frame m, video frame H corresponds to audio frame n, video frame I corresponds to audio frame o, video frame J corresponds to audio frame p, video frame K corresponds to audio frame q, and video frame L corresponds to audio frame r.
[0044] Reference Figure 6 Due to the overall delay of the video frames, video frames M, N, O, P, Q and R have no corresponding audio frames, and no sound is emitted when playing video frames M, N, O, P, Q and R.
[0045] In some embodiments, step S3 includes repairing the frequency-error video by an inverse operation of step S2 to obtain the original video.
[0046] In some embodiments, step S2 further includes encrypting the error frequency value using a preset public key to obtain an encrypted error frequency value, and encrypting and transmitting the error frequency value to prevent the error frequency value from being stolen.
[0047] In some embodiments, step S3 further includes decrypting the encrypted error frequency value using a private key corresponding to the public key to obtain the error frequency value.
[0048] In some embodiments, in step S3, the frequency-error video is repaired according to the frequency-error value.
[0049] In some embodiments, in step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frame and the video frame in the original video is adjusted according to the frequency error value to generate a frequency error video and provide it to the user. Selecting a partial time period in the original video can reduce the amount of data processing, and the generated frequency error video can also serve as a warning to the viewer, and can also reduce the viewer's discomfort to a certain extent.
[0050] Figure 7 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 7, taking the generated error frequency value of -2 as an example, audio frames a, audio frame b, audio frame c, audio frame d, audio frame e, audio frame f, audio frame g and audio frame h are delayed, and audio frame i and audio frame j are deleted. Video frame A and video frame B have no corresponding audio frames, that is, no sound is emitted when video frame A and video frame B are played; video frame C corresponds to audio frame a, video frame D corresponds to audio frame b, video frame E corresponds to audio frame c, video frame F corresponds to audio frame d, video frame G corresponds to audio frame e, video frame H corresponds to audio frame f, video frame I corresponds to audio frame g, video frame J corresponds to audio frame h, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.
[0051] Figure 8 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 8 , taking the generated error frequency value of 2 as an example, the video frame is delayed as a whole, and the audio frame k, audio frame l, audio frame m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r are delayed. Audio frame a and audio frame b have no corresponding video frames, that is, no video image is played when audio frame a and audio frame b are played. Video frame I and video frame J have no corresponding audio frames, that is, no sound is emitted when audio frame I and audio frame J are played. Video frame A corresponds to audio frame c, video frame B corresponds to audio frame d, and video frame A corresponds to audio frame c. Frame C corresponds to audio frame e, video frame D corresponds to audio frame f, video frame E corresponds to audio frame g, video frame F corresponds to audio frame h, video frame G corresponds to audio frame i, video frame H corresponds to audio frame j, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.
[0052] In some embodiments, the video platform video processing method further includes an audio recognition step, wherein the audio recognition step includes recognizing guiding words, and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address. By increasing at least one of the duration and quantity of the time period in the original video selected according to the guiding words, the discomfort of watching the wrong frequency video is continuously increased, further enhancing the warning.
[0053] In some embodiments, in step S3, after repairing the wrong video, it further includes adding annotation information to the original video obtained by repairing the wrong video to indicate that the original video obtained by repairing the wrong video is a video of a virtual character. The warning is further enhanced by adding annotation information.
[0054] In some embodiments, in step S1, the error frequency value is obtained from encryption hardware. The encryption hardware can protect the error frequency value to prevent the error frequency value from being obtained or tampered with at will.
[0055] The present invention also provides a video platform for implementing the video platform video processing method, including an acquisition unit, a frequency error adjustment unit and a frequency error adjustment recovery unit, the acquisition unit is used to obtain the frequency error value and the original video, the frequency error adjustment unit is used to adjust the timing of at least one of the audio frames and video frames in the original video according to the frequency error value to generate a frequency error video and provide it to a user, and the frequency error adjustment recovery unit is used to repair the frequency error video to obtain the original video and provide it to the user.
[0056] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.
Claims
1. A video processing method for a video platform, characterized in that: The following steps are involved: S1: Obtain pre-stored error frequency value and original video; S2: adjusting the timing of at least one of the audio frame and the video frame in the original video according to the frequency error value to generate a frequency error video and provide it to the user; S3: When the user needs to restore the video, the wrong video is repaired to obtain the original video and provide it to the user.
2. The video platform video processing method according to claim 1, characterized in that: Step S2 also includes encrypting the frequency error value using a preset public key to obtain an encrypted frequency error value.
3. The video platform video processing method according to claim 2, characterized in that: Step S3 also includes decrypting the encrypted error frequency value using a private key corresponding to the public key to obtain the error frequency value.
4. The video platform video processing method according to claim 3, characterized in that: In step S3, the frequency-error video is repaired according to the frequency-error value.
5. The video platform video processing method according to claim 1, characterized in that: In step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the original video is adjusted according to the frequency error value to generate a frequency error video and provide it to the user.
6. The video platform video processing method according to claim 5, characterized in that: The method also includes an audio recognition step, which includes recognizing guiding words and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of the guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address.
7. The video platform video processing method according to claim 1, characterized in that: In step S3, after the wrong frequency video is repaired, it also includes adding annotation information to the original video obtained by repairing the wrong frequency video to indicate that the original video obtained by repairing the wrong frequency video is a virtual character video.
8. The video platform video processing method according to claim 1, characterized in that: In step S1, the error frequency value is obtained from encryption hardware.
9. A video platform, characterized in that: The invention comprises an acquisition unit, a frequency error adjustment unit and a frequency error adjustment recovery unit, wherein the acquisition unit is used to acquire a frequency error value and an original video, the frequency error adjustment unit is used to adjust the timing of at least one of the audio frames and the video frames in the original video according to the frequency error value to generate a frequency error video and provide it to a user, and the frequency error adjustment recovery unit is used to repair the frequency error video to obtain the original video and provide it to the user.