Video processing method and video processing system

By adjusting the timing of audio frames and video frames, new videos are generated to distinguish virtual character videos, which solves the problem of malicious use of virtual human generation technology and improves the security of online interactions.

CN119946384APending Publication Date: 2025-05-06SHANGHAI TONGSHI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111534.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing virtual life generation technology may be maliciously used, resulting in ethical and security issues in online audio and video interactions, such as online fraud and virtual identity impersonation.

Method used

By obtaining random seeds to generate error frequency values, adjust the timing of audio frames and video frames in the original video, generate new videos, so that the audio frames are out of sync with video frames, thereby distinguishing virtual character videos and avoiding malicious use.

Benefits of technology

The generated new videos have discomfort and can effectively distinguish them into virtual character videos, reduce the risk of malicious use of virtual human generation technology, and improve the security of online interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946384A_ABST
    Figure CN119946384A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and system, and the method comprises the steps: obtaining a random seed, generating an error frequency value according to the random seed, adjusting the time sequence of at least one of an audio frame and a video frame in an original video according to the error frequency value, and generating a new video, according to the invention, the audio frames in the new video are not synchronized with the corresponding video frames, that is, frequency staggering exists between the video frames and the audio frames, so that the new video is uncomfortable, the new video can be distinguished as a virtual character video, and the virtual character generation technology is prevented from being maliciously used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a video processing method and a video processing system. Background Art

[0002] With the continuous advancement of virtual human generation technology, digital humans and simulated humans have been widely used in entertainment, education, medical care and other fields.

[0003] The publicly available virtual human generation technologies, such as StyleGAN face generation technology, DeepFake video generation technology, and Wav2Lip audio and video synchronization generation technology, mostly focus on the optimization of voice-driven facial expression generation algorithms, aiming to improve naturalness and synchronization, thereby improving the simulation of virtual characters.

[0004] However, this means that the current virtual human generation technology may be used maliciously, which may easily lead to ethical and security issues when conducting online audio and video interactions through virtual simulated characters, such as online fraud and virtual identity impersonation.

[0005] Therefore, it is necessary to provide a new video processing method and a video processing system to solve the above problems existing in the prior art. Summary of the invention

[0006] The purpose of the present invention is to provide a video processing method and a video processing system, which can distinguish a video as a virtual character video, thereby preventing the virtual human generation technology from being maliciously used.

[0007] To achieve the above object, the video processing method of the present invention comprises the following steps:

[0008] S1: Obtain a random seed, and then generate an error frequency value according to the random seed;

[0009] S2: adjusting the timing of at least one of the audio frames and the video frames in the original video according to the frequency error value to generate a new video.

[0010] The beneficial effect of the video processing method is: obtaining a random seed, then generating a frequency error value according to the random seed, and adjusting the timing of at least one of the audio frames and video frames in the original video according to the frequency error value to generate a new video, so that the audio frame in the new video is not synchronized with the corresponding video frame, that is, there is a frequency error between the video frame and the audio frame, so that the new video has an uncomfortable feeling and can be distinguished as a virtual character video, thereby avoiding the malicious use of virtual human generation technology.

[0011] Optionally, in step S1, a random seed is obtained from encryption hardware.

[0012] Optionally, the video processing method further includes a seed safety detection step, and the seed safety detection step includes:

[0013] Signing the seed with a preset private key to obtain a digital signature;

[0014] The digital signature is verified by means of a public key.

[0015] Optionally, in step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the partial time period is adjusted according to the frequency error value to generate a new video.

[0016] Optionally, before executing step S1, it also includes selecting a significant frequency error level or a subtle frequency error level. When executing step S1, a random seed is obtained according to the selected significant frequency error level or the subtle frequency error level, wherein the frequency error value generated by the random seed obtained according to the significant frequency error level is greater than the frequency error value generated by the random seed obtained according to the subtle frequency error level.

[0017] Optionally, after selecting the significant error frequency level, an audio recognition step is also included, the audio recognition step includes identifying guiding words, and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of the guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address.

[0018] Optionally, in step S1, generating the error frequency value according to the random seed includes: converting the random seed into the error frequency value through the formula X_{n+1}=(aX_n+c)mod m, wherein X_n is the random seed, X_{n+1} is the error frequency value, and a, c and m are preset constants.

[0019] Optionally, step S2 further includes selecting a random frequency error algorithm, and then adjusting the timing of at least one of the audio frames and video frames in the original video according to the frequency error value through the selected random frequency error algorithm to generate a new video.

[0020] The present invention also provides a video processing system, including a seed acquisition unit, a frequency error value calculation unit and a frequency error adjustment unit, wherein the seed acquisition unit is used to acquire a random seed, the frequency error value calculation unit is used to generate a frequency error value according to the random seed, and the frequency error adjustment unit is used to adjust the timing of at least one of the audio frames and video frames in the original video according to the frequency error value to generate a new video.

[0021] The beneficial effect of the video processing system is that: the seed acquisition unit acquires a random seed, the frequency error value calculation unit generates a frequency error value according to the random seed, and the frequency error adjustment unit adjusts the timing of at least one of the audio frame and the video frame in the original video according to the frequency error value to generate a new video, so that the audio frame in the new video is not synchronized with the corresponding video frame, that is, there is a frequency error between the video frame and the audio frame, so that the new video has an uncomfortable feeling and can be distinguished as a virtual character video, thereby avoiding the malicious use of virtual human generation technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flowchart of a video processing method in some embodiments of the present invention;

[0023] Figure 2 A schematic diagram of the timing of coordination of audio frames and video frames in an original video in some embodiments of the present invention;

[0024] Figure 3 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;

[0025] Figure 4 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;

[0026] Figure 5 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;

[0027] Figure 6 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;

[0028] Figure 7 In some embodiments of the present invention, Figure 2 A schematic diagram of the timing of the coordination of audio frames and video frames in a new video generated from an original video is shown;

[0029] Figure 8 In some embodiments of the present invention, Figure 2 Schematic diagram of the coordination timing of audio frames and video frames in the new video generated from the original video. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0031] In view of the problems existing in the prior art, an embodiment of the present invention provides a video processing method. Figure 1 , the video processing method comprises the following steps:

[0032] S1: Obtain a random seed, and then generate an error frequency value according to the random seed;

[0033] S2: adjusting the timing of at least one of the audio frames and the video frames in the original video according to the frequency error value to generate a new video.

[0034] In some embodiments, when executing step S1, the same or different error frequency values ​​can be continuously obtained, so that when executing step S2, the same or different timing adjustments can be made to different time periods of the audio frames and video frames in the original video.

[0035] In some embodiments, in step S1, a random seed is obtained from the encryption hardware. In some embodiments, the video processing method further includes a seed security detection step, and the seed security detection step includes: signing the seed with a preset private key to obtain a digital signature; and verifying the digital signature with a public key. A number of the seeds are pre-stored in the encryption hardware, and the encryption hardware ensures that the seeds cannot be obtained at will through the encryption and decryption algorithm, and through the seed security detection step, it can be ensured that the seeds will not be tampered with at will, thereby ensuring the integrity of the seeds and will not be forged, thereby ensuring that the subsequently generated frequency error values ​​are within the valid range, and will not cause the timing difference between the audio frame and the video frame in the generated new video after executing step S2 to be too large, so that the new video can be perceived as unnatural by humans, but it does not seriously affect people's viewing of the new video.

[0036] Figure 2 FIG. 1 is a schematic diagram of the timing of the coordination of audio frames and video frames in the original video in some embodiments of the present invention. Figure 2, the original video includes video frame A, video frame B, video frame C, video frame D, video frame E, video frame F, video frame G, video frame H, video frame I, video frame J, video frame K, video frame L, video frame M, video frame N, video frame O, video frame P, video frame Q, video frame R, audio frame a, audio frame b, audio frame c, audio frame d, audio frame e, audio frame f, audio frame g, audio frame h, audio frame i, audio frame j, audio frame k, audio frame l, audio frame m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r.

[0037] Reference Figure 2 , video frame A corresponds to audio frame a, video frame B corresponds to audio frame b, video frame C corresponds to audio frame c, video frame D corresponds to audio frame d, video frame E corresponds to audio frame e, video frame F corresponds to audio frame f, video frame G corresponds to audio frame g, video frame H corresponds to audio frame h, video frame I corresponds to audio frame i, video frame J corresponds to audio frame j, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.

[0038] Figure 3 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 3 , taking the generated error frequency value of -2 as an example, the audio frame is delayed as a whole, and video frame A and video frame B have no corresponding audio frames, that is, no sound is emitted when video frame A and video frame B are played; video frame C corresponds to audio frame a, video frame D corresponds to audio frame b, video frame E corresponds to audio frame c, video frame F corresponds to audio frame d, video frame G corresponds to audio frame e, video frame H corresponds to audio frame f, video frame I corresponds to audio frame g, video frame J corresponds to audio frame h, video frame K corresponds to audio frame i, video frame L corresponds to audio frame j, video frame M corresponds to audio frame k, video frame N corresponds to audio frame l, video frame O corresponds to audio frame m, video frame P corresponds to audio frame n, video frame Q corresponds to audio frame o, and video frame R corresponds to audio frame p.

[0039] Reference Figure 3 Due to the overall delay of the audio frames, the audio frames q and r have no corresponding video frames. The audio frames q and r can be played without video images, or the audio frames q and r can be deleted.

[0040] Figure 4 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 4 , taking the generated error frequency value of -6 as an example, the audio frame is delayed as a whole, and video frame A, video frame B, video frame C, video frame D, video frame E and video frame F have no corresponding audio frames, that is, no sound is emitted when playing video frame A, video frame B, video frame C, video frame D, video frame E and video frame F; video frame G corresponds to audio frame a, video frame H corresponds to audio frame b, video frame I corresponds to audio frame c, video frame J corresponds to audio frame d, video frame K corresponds to audio frame e, video frame L corresponds to audio frame f, video frame M corresponds to audio frame g, video frame N corresponds to audio frame h, video frame O corresponds to audio frame i, video frame P corresponds to audio frame j, video frame Q corresponds to audio frame k, and video frame R corresponds to audio frame l.

[0041] Reference Figure 4 Due to the overall delay of audio frames, audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r have no corresponding video frames. Audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r can be played without video images, and audio frames m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r can also be deleted.

[0042] Figure 5 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 5 , taking the generated error frequency value of 2 as an example, the video frame is delayed as a whole, and audio frame a and audio frame b have no corresponding video frames, that is, no video image is played when audio frame a and audio frame b are played; video frame A corresponds to audio frame c, video frame B corresponds to audio frame d, video frame C corresponds to audio frame e, video frame D corresponds to audio frame f, video frame E corresponds to audio frame g, video frame F corresponds to audio frame h, video frame G corresponds to audio frame i, video frame H corresponds to audio frame j, video frame I corresponds to audio frame k, video frame J corresponds to audio frame l, video frame K corresponds to audio frame m, video frame L corresponds to audio frame n, video frame M corresponds to audio frame o, video frame N corresponds to audio frame p, video frame O corresponds to audio frame q, and video frame P corresponds to audio frame r.

[0043] Reference Figure 5 Due to the overall delay of the video frames, video frames Q and R have no corresponding audio frames, and no sound is emitted when playing video frames Q and R.

[0044] Figure 6 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 6, taking the generated error frequency value of 6 as an example, the video frame is delayed as a whole, and audio frame a, audio frame b, audio frame c, audio frame d, audio frame e and audio frame f have no corresponding video frames, that is, no video image is played when audio frame a, audio frame b, audio frame c, audio frame d, audio frame e and audio frame f are played; video frame A corresponds to audio frame g, video frame B corresponds to audio frame h, video frame C corresponds to audio frame i, video frame D corresponds to audio frame j, video frame E corresponds to audio frame k, video frame F corresponds to audio frame l, video frame G corresponds to audio frame m, video frame H corresponds to audio frame n, video frame I corresponds to audio frame o, video frame J corresponds to audio frame p, video frame K corresponds to audio frame q, and video frame L corresponds to audio frame r.

[0045] Reference Figure 6 Due to the overall delay of the video frames, video frames M, N, O, P, Q and R have no corresponding audio frames, and no sound is emitted when playing video frames M, N, O, P, Q and R.

[0046] In some embodiments, in step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the partial time period is adjusted according to the error frequency value to generate a new video. Selecting a partial time period in the original video can reduce the amount of data processing, and the generated new video can also serve as a warning to the viewer, and can also reduce the discomfort of the viewer to a certain extent.

[0047] Figure 7 In some embodiments of the present invention, Figure 2 The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 7 , taking the generated error frequency value of -2 as an example, audio frames a, audio frame b, audio frame c, audio frame d, audio frame e, audio frame f, audio frame g and audio frame h are delayed, and audio frame i and audio frame j are deleted. Video frame A and video frame B have no corresponding audio frames, that is, no sound is emitted when video frame A and video frame B are played; video frame C corresponds to audio frame a, video frame D corresponds to audio frame b, video frame E corresponds to audio frame c, video frame F corresponds to audio frame d, video frame G corresponds to audio frame e, video frame H corresponds to audio frame f, video frame I corresponds to audio frame g, video frame J corresponds to audio frame h, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.

[0048] Figure 8 In some embodiments of the present invention, Figure 2The following is a schematic diagram of the timing of the audio and video frames in the new video generated from the original video. Figure 8 , taking the generated error frequency value of 2 as an example, the video frame is delayed as a whole, and the audio frame k, audio frame l, audio frame m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r are delayed. Audio frame a and audio frame b have no corresponding video frames, that is, no video image is played when audio frame a and audio frame b are played. Video frame I and video frame J have no corresponding audio frames, that is, no sound is emitted when audio frame I and audio frame J are played. Video frame A corresponds to audio frame c, video frame B corresponds to audio frame d, and video frame A corresponds to audio frame c. Frame C corresponds to audio frame e, video frame D corresponds to audio frame f, video frame E corresponds to audio frame g, video frame F corresponds to audio frame h, video frame G corresponds to audio frame i, video frame H corresponds to audio frame j, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.

[0049] In some embodiments, before executing step S1, it also includes selecting a significant frequency error level or a subtle frequency error level. When executing step S1, a random seed is obtained according to the selected significant frequency error level or the subtle frequency error level, wherein the frequency error value generated by the random seed obtained according to the significant frequency error level is greater than the frequency error value generated by the random seed obtained according to the subtle frequency error level. Specifically, after selecting the significant frequency error level, when watching the new video, the viewer can feel that all or part of the video frames and audio frames do not correspond obviously, and there is a sense of discomfort in watching; after selecting the subtle frequency error level, when watching the new video, the viewer can hardly feel that all or part of the video frames and audio frames do not correspond obviously, and there is no sense of discomfort in watching, and it can only be detected by machine.

[0050] In some embodiments, after selecting the significant frequency error level, an audio recognition step is further included, the audio recognition step includes recognizing guiding words, and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address. By increasing at least one of the duration and quantity of the time period in the original video selected according to the guiding words, the discomfort of watching the new video is continuously increased, further enhancing the warning.

[0051] In some embodiments, in step S1, generating the frequency error value according to the random seed includes: converting the random seed into the frequency error value through the formula X_{n+1}=(aX_n+c)mod m, wherein X_n is the random seed, X_{n+1} is the frequency error value, and a, c and m are preset constants.

[0052] In some embodiments, step S2 further includes selecting a random frequency error algorithm, and then adjusting the timing of at least one of the audio frame and the video frame in the original video according to the frequency error value by using the selected random frequency error algorithm to generate a new video. Different frequency error algorithms have certain differences in the methods of adjusting the original video, which increases the difficulty of restoring the new video to the original video and also increases the difficulty of adapting the original video to the frequency error algorithm.

[0053] The present invention also provides a video processing system for implementing the video processing method, the video processing system includes a seed acquisition unit, a frequency error value calculation unit and a frequency error adjustment unit, the seed acquisition unit is used to obtain a random seed, the frequency error value calculation unit is used to generate a frequency error value according to the random seed, and the frequency error adjustment unit is used to adjust the timing of at least one of the audio frames and video frames in the original video according to the frequency error value to generate a new video.

[0054] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.

Claims

1. A video processing method, characterized in that: The following steps are involved: S1: Obtain a random seed, and then generate an error frequency value according to the random seed; S2: adjusting the timing of at least one of the audio frames and the video frames in the original video according to the frequency error value to generate a new video.

2. The video processing method according to claim 1, characterized in that: In step S1, a random seed is obtained from encryption hardware.

3. The video processing method according to claim 1 or 2, characterized in that: The invention also includes a seed safety detection step, wherein the seed safety detection step includes: Signing the seed with a preset private key to obtain a digital signature; The digital signature is verified by means of a public key.

4. The video processing method according to claim 1, characterized in that: In step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the partial time period is adjusted according to the frequency error value to generate a new video.

5. The video processing method according to claim 1 or 2, characterized in that: Before executing step S1, it also includes selecting a significant frequency error level or a subtle frequency error level. When executing step S1, a random seed is obtained according to the selected significant frequency error level or the subtle frequency error level, wherein the frequency error value generated by the random seed obtained according to the significant frequency error level is greater than the frequency error value generated by the random seed obtained according to the subtle frequency error level.

6. The video processing method according to claim 5, characterized in that: In step S2, a partial time period in the original video is selected, and the timing of at least one of the audio frames and the video frames in the partial time period is adjusted according to the frequency error value to generate a new video.

7. The video processing method according to claim 6, characterized in that: After selecting the significant frequency error level, an audio recognition step is also included, wherein the audio recognition step includes identifying guiding words, and increasing at least one of the duration and quantity of the time period in the original video selected in step S2 as the number of the guiding words increases, wherein the guiding words include remittance, transfer, account number, password, identity, and address.

8. The video processing method according to claim 1, characterized in that: In step S1, generating the frequency error value according to the random seed includes: converting the random seed into the frequency error value by the formula X_{n+1}=(aX_n+c)modm, wherein X_n is the random seed, X_{n+1} is the frequency error value, and a, c and m are preset constants.

9. The video processing method according to claim 1, characterized in that: Step S2 also includes selecting a random frequency error algorithm, and then adjusting the timing of at least one of the audio frames and video frames in the original video according to the frequency error value through the selected random frequency error algorithm to generate a new video.

10. A video processing system, characterized in that: The invention comprises a seed acquisition unit, a frequency error value calculation unit and a frequency error adjustment unit, wherein the seed acquisition unit is used to acquire a random seed, the frequency error value calculation unit is used to generate a frequency error value according to the random seed, and the frequency error adjustment unit is used to adjust the timing of at least one of the audio frame and the video frame in the original video according to the frequency error value to generate a new video.