Video processing method and device, computer device and storage medium

By adding image perturbation noise and audio interference information to video processing, combined with noise reduction processing, the problem of reduced effective information in traditional methods is solved, thus achieving secure protection of user information.

CN115760634BActive Publication Date: 2026-02-27INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211507006.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-02-27
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

Traditional video processing methods reduce the amount of effective information in videos, failing to effectively protect user information security.

Method used

By adding image perturbation noise and audio interference information through image edge recognition strategy, combined with noise reduction processing, target information in the video is extracted and processed.

Benefits of technology

It increases the amount of effective information after video processing and effectively protects user information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760634B_ABST
    Figure CN115760634B_ABST
Patent Text Reader

Abstract

The application relates to a video processing method and device, computer equipment and a storage medium. The application relates to the field of artificial intelligence technology, and the method comprises the following steps: acquiring target video information; extracting feature images of each single-frame picture information, and identifying target image edge information in each feature image through a picture edge identification strategy; adding image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining the processed single-frame picture information into processed picture information; adding audio interference information to audio information, and performing sound elimination processing on target audio data in the audio information that meets a sound elimination processing condition to obtain processed audio information; and determining processed video information according to the processed audio information and the processed picture information. The method can improve the retention degree of effective information of the processed video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a video processing method and device, computer equipment and storage medium. BACKGROUND

[0002] With the online trend of the financial industry, more and more financial services in the form of audio and video calls are transferred from offline to online, and video face review is one of the most widely used scenarios. In this scenario, a large amount of unstructured face review video data will be generated, and such data has a wide range of uses, such as service quality inspection, or emotion recognition in financial scenarios, anti-fraud research, etc. However, due to the presence of user information that needs to be protected in the video, such as user face, voiceprint, and other biological features, property certificate, ID card, and other certificate information, and spoken phone numbers, home addresses, and other personal information, it cannot be directly used and needs to be processed.

[0003] The traditional technical solution usually determines the target image content to be processed in the video first, then formulates a processing rule (for example, processing the license plate in the image), inputs each frame of image of the video to be processed into a target neural network, and then performs blurring processing on the recognized target object; and the audio in the video to be processed is muted, thereby completing the video processing process. However, too much blurring processing of the picture information content of the video and muting of the audio information will result in less effective information in the processed video. SUMMARY

[0004] Therefore, it is necessary to provide a video processing method, device, computer equipment, computer readable storage medium, and computer program product to solve the above technical problems.

[0005] In a first aspect, the present application provides a video processing method. The method comprises:

[0006] obtaining target video information; the video information comprises a plurality of single-frame picture information and audio information;

[0007] extracting a feature image of each single-frame picture information, and identifying target image edge information in each feature image through a picture edge identification strategy;

[0008] adding image disturbance noise to the image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining each processed single-frame picture information into processed picture information; the number of processed single-frame picture information is the same as the number of single-frame picture information;

[0009] Adding audio interference information in the audio information, and performing mute processing on target audio data in the audio information satisfying a mute processing condition to obtain processed audio information;

[0010] According to the processed audio information and the processed picture information, determine the processed video information.

[0011] Optionally, the feature image of each single-frame picture information includes:

[0012] Frame processing the picture information to obtain a plurality of single-frame picture information;

[0013] For each single-frame picture information, the single-frame picture information is divided into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information.

[0014] Extract the initial feature image of each picture information of adjacent single-frame picture information in time sequence, and integrate the initial feature image of adjacent picture information according to time sequence, and extract the dynamic interaction feature of the integrated initial feature image to obtain the feature image of adjacent single-frame picture information.

[0015] Optionally, the target image edge information in each feature image is identified by a picture edge recognition strategy, including:

[0016] For each feature image, the target image information in the feature image is identified by a self-attention recognition strategy, and the text information of the feature image is identified by a text recognition algorithm; the target image information is the image information of the target image in the feature image.

[0017] According to the text information of the feature image and the target image information, determine the target image edge information corresponding to the feature image.

[0018] Optionally, the image disturbance noise is added to the image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and each processed single-frame picture information is combined into processed picture information, including:

[0019] Adding image disturbance noise to the target image information in the target image edge information of each single-frame picture information, and performing blurring processing on the text information in the target image edge information to obtain the processed single-frame picture information corresponding to the single-frame picture information.

[0020] All processed single-frame picture information is combined into processed picture information according to the time sequence of the picture information.

[0021] Optionally, the adding audio interference information in the audio information comprises:

[0022] performing speech recognition processing on the audio information to obtain audio text information corresponding to each audio data in the audio information, and determining whether there is user audio text information in the audio text information;

[0023] in a case where there is user audio text information in the audio text information, adding audio interference information in audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting user voiceprint features of the audio data.

[0024] Optionally, the muting processing on the target audio data in the audio information to obtain the processed audio information comprises:

[0025] extracting target text information of the audio text information according to a preset target information screening strategy, and taking audio data corresponding to the target text information as target audio data;

[0026] performing muting processing on the target audio data in the initial processed audio information to obtain the processed audio information.

[0027] In a second aspect, the present application further provides a video processing device. The device comprises:

[0028] an acquisition module configured to acquire target video information; the video information comprises a plurality of single-frame picture information and audio information;

[0029] an extraction module configured to extract feature images of each single-frame picture information, and identify target image edge information in each feature image through a picture edge identification strategy;

[0030] a processing module configured to add image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combine each processed single-frame picture information into processed picture information; the number of the processed single-frame picture information is the same as the number of the single-frame picture information;

[0031] a muting module configured to add audio interference information in the audio information, and perform muting processing on target audio data in the audio information that meets a muting processing condition to obtain processed audio information;

[0032] a determination module configured to determine processed video information according to the processed audio information and the processed picture information.

[0033] Optionally, the extraction module is specifically configured to:

[0034] frame the picture information to obtain a plurality of single-frame picture information;

[0035] For each single-frame picture information, the single-frame picture information is divided into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information.

[0036] extract the initial feature image of each picture information of adjacent single-frame picture information in the time sequence, integrate the initial feature images of adjacent picture information according to the time sequence, and extract the dynamic interaction feature of the integrated initial feature image to obtain the feature image of adjacent single-frame picture information.

[0037] Optionally, the extraction module is specifically used for:

[0038] For each feature image, the target image information in the feature image is identified through a self-attention recognition strategy, and the text information of the feature image is identified through a text recognition algorithm; the target image information is the image information of the target image in the feature image.

[0039] According to the text information of the feature image and the target image information, the target image edge information corresponding to the feature image is determined.

[0040] Optionally, the processing module is specifically used for:

[0041] In the target image information in the target image edge information of each single-frame picture information, image disturbance noise is added, and the text information in the target image edge information is blurred to obtain the processed single-frame picture information corresponding to the single-frame picture information.

[0042] All processed single-frame picture information is combined into processed picture information according to the time sequence of the picture information.

[0043] Optionally, the sound elimination module is specifically used for:

[0044] The audio information is subjected to speech recognition processing to obtain audio text information corresponding to each audio data in the audio information, and it is judged whether there is user audio text information in the audio text information.

[0045] In the case that there is user audio text information in the audio text information, audio interference information is added in the audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting the user voiceprint feature of the audio data.

[0046] Optionally, the sound elimination module is specifically used for:

[0047] According to the preset target information filtering strategy, the target text information of the audio text information is extracted, and the audio data corresponding to the target text information is used as the target audio data;

[0048] The target audio data in the initially processed audio information is muted to obtain the processed audio information.

[0049] Thirdly, this application provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in any one of the first aspects.

[0050] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any one of the first aspects.

[0051] Fifthly, this application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.

[0052] The aforementioned video processing method, apparatus, computer equipment, and storage medium acquire video information to be reviewed; the video information includes image information and audio information; divide the image information into multiple single-frame image information, extract feature images of each single-frame image information, and identify target image edge information in each single-frame image information through an image edge recognition strategy; add image perturbation noise to the image information corresponding to each target image edge information to obtain multiple processed single-frame image information, and combine the processed single-frame image information into processed image information; the number of processed single-frame image information is the same as the number of single-frame image information; add audio interference information to the audio information, and perform noise reduction processing on the target audio data in the audio information to obtain processed audio information; determine the processed video information based on the processed audio information and the processed image information. By adding noise and partially blurring the target information in the image information, the retention effect of effective information in the processed image information is improved, and by adding interference and noise reduction processing to the target information in the audio information, the amount of effective information contained in the processed video is increased. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating a video processing method in one embodiment;

[0054] Figure 2A flowchart of a process for determining target image edge information in an embodiment;

[0055] Figure 3 A flowchart of a process for adding audio interference information in an embodiment;

[0056] Figure 4 A flowchart of a process for video processing in an embodiment;

[0057] Figure 5 A block diagram of a structure of a video processing device in an embodiment;

[0058] Figure 6 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0059] To make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0060] The video processing method provided by the embodiments of the present application can be applied to a terminal, a server, a system including a terminal and a server, and realized through the interaction of the terminal and the server. The terminal can include, but is not limited to, various personal computers, notebook computers, tablet computers, etc. The terminal adds noise and partial blurring processing to the target information in the picture information, improves the retention effect of the effective information of the processed picture information, adds interference and sound elimination processing to the target information in the audio information, improves the readability and the processed effect of the processed audio information, and thus improves the retention degree of the effective information of the processed video.

[0061] In an embodiment, as shown in Figure 1 A video processing and identification method is provided. The method is applied to a terminal as an example and includes the following steps:

[0062] Step S101, obtaining target video information.

[0063] The video information includes multiple single-frame picture information and audio information.

[0064] In the embodiment, the terminal receives video information to be processed transmitted by each channel through multiple video information transmission channels, and divides the multiple video information into multiple to-be-reviewed video information according to the order of video transmission. The video information can be, but is not limited to, monitoring video, live video, recorded video, shooting video, and other data information combining picture and audio.

[0065] Step S102, extract the feature image of each single-frame picture information, and identify the target image edge information in each feature image through a picture edge identification strategy.

[0066] In this embodiment, the terminal divides the picture information in the video information into multiple single-frame picture information according to the frame number of the picture information, and extracts the feature image in each single-frame picture information through a feature image extraction network. The feature image extraction network can be, but is not limited to, any neural network that can implement the above steps. The terminal locates the target image in each feature image through a picture edge identification strategy, and identifies the edge information of each target image. The edge information of the target image is the image edge contour information of the target image. For example, the target image is a portrait, and the edge information of the target image is the edge contour information of the face. The feature image is a plurality of images containing valid content information in the single-frame picture information. The feature image can be, but is not limited to, a human image, a plant image, an article image, a text image, and the like, which can represent the image information in the single-frame picture information. The feature image extraction network can be, but is not limited to, any network that can implement the above feature image extraction step, such as a deep learning neural network, a convolutional neural network, and the like.

[0067] Step S103, adding image disturbance noise to the image information corresponding to each target image edge information to obtain multiple processed single-frame picture information, and combining each processed single-frame picture information into processed picture information.

[0068] The number of processed single-frame picture information is the same as the number of single-frame picture information.

[0069] In this embodiment, the terminal adds image disturbance noise to the image information corresponding to each target edge information, thereby processing the image information to obtain processed image information (i.e., processed single-frame picture information). The image disturbance noise can be, but is not limited to, an image adversarial sample. The image adversarial sample can make the image information lose the feature image that can be identified by the identification system. For example, the image information is a face image, and the processed image information is an image information that loses the original biological feature that can be input into the face recognition system.

[0070] The terminal arranges all the processed single-frame picture information in the time order of the frame numbers to obtain the processed picture information.

[0071] Step S104, adding audio interference information to the audio information, and performing mute processing on the target audio data in the audio information that meets the mute processing condition to obtain processed audio information.

[0072] In this embodiment, the terminal adds audio interference information to the audio information, and identifies target audio data in the audio information. The terminal performs noise reduction processing on the target audio data in the audio information to obtain processed audio information. The specific audio processing process will be described in detail later.

[0073] In step S105, the processed video information is determined according to the processed audio information and the processed picture information.

[0074] In this embodiment, the terminal splices the processed audio information and the processed picture information to obtain the processed video information.

[0075] Based on the above scheme, by adding noise and partial blurring processing to the target information in the picture information, the retention effect of the effective information of the processed picture information is improved, and by adding interference and noise reduction processing to the target information in the audio information, the information amount of the effective information contained in the processed video is improved.

[0076] Optionally, the feature image of each single-frame picture information is extracted, including: performing frame processing on the picture information to obtain a plurality of single-frame picture information; for each single-frame picture information, the single-frame picture information is segmented into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information; the initial feature images of the picture information of adjacent single-frame picture information in time sequence are extracted, and the initial feature images of adjacent picture information are integrated according to time sequence, and the dynamic interaction features of the integrated initial feature images are extracted to obtain the feature images of adjacent single-frame picture information.

[0077] In this embodiment, the terminal first performs frame processing on the picture information according to the frame number to obtain a plurality of single-frame picture information. For each single-frame picture information, the terminal segments the single-frame picture information into a plurality of fixed-size images, and compresses the dimension of each image by linear projection to reduce the dimension, to obtain a plurality of picture information. The terminal extracts the feature image (i.e. initial feature image) of each picture information, and integrates the initial feature images of the picture information at the same position of adjacent single-frame picture information in a sub-time sequence according to the time sequence, and then identifies the dynamic features contained in the picture information based on all the integrated initial feature images. For example, the dynamic features can represent the stretching process of the person in the picture and the bending process of the person. The terminal takes the feature image of each picture information in the single-frame picture information and the dynamic features as the feature image of the picture information, and takes the feature images of all picture information in the single-frame picture information as the feature image of the single-frame picture information.

[0078] Based on the above scheme, by segmenting the single-frame picture information to extract the feature image in the single-frame picture information, the accuracy of the extracted feature image in the single-frame picture information is improved.

[0079] Optional, such as Figure 2 As shown, the target image edge information in each feature image is identified through an image edge recognition strategy, including:

[0080] Step S201: For each feature image, the target image information in the feature image is identified through a self-attention recognition strategy, and the text information in the feature image is identified through a text recognition algorithm.

[0081] Among them, the target image information is the image information of the target image in the feature image.

[0082] Step S202: Determine the edge information of the target image corresponding to the feature image based on the text information of the feature image and the target image information.

[0083] In this embodiment, as Figure 2 As shown, for each feature image, the terminal uses a Transformer self-attention strategy to identify the target image (i.e., target image information) within that feature image, and then uses ORC (Optical Character Recognition) technology to identify the text information within that feature image. Next, the terminal uses decoder technology to perform collaborative localization processing on the identified target images and text information, locating the edge information of each target image and text information into the single-frame image information corresponding to that feature image, thus obtaining the target image edge information corresponding to that feature image. The target image can be, but is not limited to, images of faces, identification documents, objects, or other images containing information that the user needs to protect. The text information can be, but is not limited to, numerical information, text information, character information, such as names, phone numbers, email addresses, addresses, ID numbers, and other text information involving personal information that needs to be protected.

[0084] Based on the above scheme, by locating the target image in the feature image and the text information in the feature image, the accuracy of the processed feature image is improved, and the effective information of the single frame image is preserved to the greatest extent.

[0085] Optionally, image perturbation noise is added to the image information corresponding to the edge information of each target image to obtain multiple processed single-frame image information, and the processed single-frame image information is combined into processed image information, including: adding image perturbation noise to the target image information in the target image edge information of each single-frame image information, and blurring the text information in the target image edge information to obtain the processed single-frame image information corresponding to the single-frame image information; and combining all the processed single-frame image information into processed image information according to the time sequence of the image information.

[0086] In this embodiment, the terminal adds picture image disturbance noise to the target image information in the target image edge information of each single-frame picture information, so as to eliminate the characteristic image of the image information, wherein the picture image disturbance noise can be but is not limited to any anti-sample that can achieve the above purpose. The terminal blurs the text information in the target image edge information of the single-frame picture information, so as to eliminate the distinguishability of the text information. The terminal takes the single-frame picture information whose characteristic image has been eliminated and whose text distinguishability has been eliminated as the processed single-frame picture information. The way of adding image disturbance noise and the way of blurring text information can be but are not limited to any way that can achieve the above steps.

[0087] The terminal combines all the processed single-frame picture information according to the time arrangement order of each single-frame picture information to obtain the processed picture information.

[0088] Based on the above scheme, by adding picture image disturbance noise to the target image information and blurring only the text information, the processed single-frame picture information is obtained, which improves the retention degree of effective information of the processed picture information.

[0089] Optionally, as shown in Figure 3 adding audio interference information in the audio information, including:

[0090] Step S301, performing speech recognition processing on the audio information to obtain audio text information corresponding to each audio data in the audio information, and determining whether there is user audio text information in the audio text information.

[0091] Step S302, in the case that there is user audio text information in the audio text information, adding audio interference information in the audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting the user voiceprint feature of the audio data.

[0092] In this embodiment, the terminal performs speech recognition processing on the audio information and converts the audio information into audio text information corresponding to the audio information. The terminal establishes a correspondence between each piece of text information in the audio text information and each audio period in the audio information, and adds a timestamp at the start time point and the end time point of each audio period to obtain audio data corresponding to each piece of text information. The terminal identifies whether there is user audio text information in the audio text information by using a text recognition technology. If there is no user audio text information in the audio text information, the terminal proceeds to the next step. If there is user audio text information in the audio text information, the terminal determines the text information included in the user text audio information and queries the user audio data corresponding to each piece of text information (i.e., the audio data corresponding to the user audio text information) by using the correspondence between each piece of text information and the audio data. The terminal adds audio interference information to each piece of user audio data to obtain initial processed audio information. The audio interference information can be audio adversarial samples, audio noise information, or other interference information that can eliminate user voice feature information.

[0093] Based on the above scheme, by eliminating the user voice feature, the retention degree of the effective information of the processed audio information is further improved on the basis of improving the protection degree of the user information to be protected in the processed audio information

[0094] Optionally, the target audio data in the audio information is de-voiced to obtain the processed audio information, including: extracting target text information of the audio text information according to a preset target information screening strategy, and taking the audio data corresponding to the target text information as the target audio data; and de-voicing the target audio data in the initial processed audio information to obtain the processed audio information.

[0095] In this embodiment, based on the initial processed audio information, the terminal predefines a target text data recognition strategy, and identifies target text information in the text information and determines target text information included in each target text information by using a text recognition technology according to the target text data recognition strategy. The terminal queries the target audio data corresponding to each target text information in the initial processed audio information by using the correspondence between each piece of text information and the audio data, and de-voices each target audio data to obtain the processed audio information.

[0096] Based on the above scheme, by de-voicing the target audio data corresponding to the target information, the protection degree of the user privacy in the processed audio information is improved.

[0097] The application also provides a video processing example, as shown in Figure 4 The specific processing process includes the following steps:

[0098] Step S401, obtain target video information.

[0099] Step S402, frame information is processed to obtain a plurality of single frame picture information.

[0100] Step S403, for each single frame picture information, the single frame picture information is divided into a plurality of picture information, and a plurality of picture information of each single frame picture information is obtained.

[0101] Step S404, the initial feature image of each picture information of adjacent single frame picture information in time sequence is extracted, and the initial feature images of adjacent picture information are integrated according to time sequence, and the dynamic interaction feature of the integrated initial feature image is extracted, and the feature image of adjacent single frame picture information is obtained.

[0102] Step S405, for each feature image, the target image information in the feature image is identified by self-attention recognition strategy, and the text information of the feature image is identified by text recognition algorithm.

[0103] Step S406, according to the text information and target image information of the feature image, the target image edge information corresponding to the feature image is determined.

[0104] Step S407, adding image disturbance noise to the target image information in the target image edge information of each single frame picture information, and blurring the text information in the target image edge information, to obtain the processed single frame picture information corresponding to the single frame picture information.

[0105] Step S408, all processed single frame picture information is combined into processed picture information according to the time sequence of picture information.

[0106] Step S409, the audio information is processed by speech recognition to obtain audio text information corresponding to each audio data in the audio information, and it is judged whether there is user audio text information in the audio text information.

[0107] Step S410, in the case that there is user audio text information in the audio text information, audio interference information is added in the audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting the user voiceprint feature of the audio data.

[0108] Step S411, according to the preset target information screening strategy, the target text information of the audio text information is extracted, and the audio data corresponding to the target text information is taken as the target audio data.

[0109] Step S412, the target audio data in the initial processed audio information is processed to obtain the processed audio information.

[0110] In step S413, the processed video information is determined according to the processed audio information and the processed picture information.

[0111] It should be understood that, although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0112] Based on the same inventive concept, the embodiments of the present application also provide a video processing apparatus for implementing the above-mentioned video processing method. The problem-solving implementation scheme provided by the apparatus is similar to the implementation scheme described in the above method, so the specific limitations in one or more video processing apparatus embodiments provided below can refer to the limitations of the video processing method described above, which will not be repeated here.

[0113] In one embodiment, as shown in FIG. 5, a video processing apparatus is provided, which includes an acquisition module 510, an extraction module 520, a processing module 530, a de-voice module 540, and a determination module 550, wherein: Figure 5 The acquisition module 510 is configured to acquire target video information; the video information includes a plurality of single-frame picture information and audio information.

[0114] The extraction module 520 is configured to extract a feature image of each single-frame picture information, and identify target image edge information in each feature image through a picture edge identification strategy.

[0115] The processing module 530 is configured to add image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combine each processed single-frame picture information into processed picture information; the number of processed single-frame picture information is the same as the number of single-frame picture information.

[0116] The de-voice module 540 is configured to add audio interference information in the audio information, and perform de-voice processing on target audio data in the audio information to obtain processed audio information.

[0117] The determination module 550 is configured to determine processed video information according to the processed audio information and the processed picture information.

[0118] determining module 550 is configured to determine processed video information according to the processed audio information and the processed picture information.

[0119] Optionally, the extraction module 520 is specifically configured to:

[0120] frame processing the picture information to obtain a plurality of single-frame picture information;

[0121] for each single-frame picture information, segmenting the single-frame picture information into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information;

[0122] extracting initial feature images of the picture information of adjacent single-frame picture information in a time sequence, integrating the initial feature images of adjacent picture information according to the time sequence, and extracting dynamic interaction features of the integrated initial feature images to obtain feature images of adjacent single-frame picture information.

[0123] Optionally, the extraction module 520 is specifically configured to:

[0124] for each feature image, identifying target image information in the feature image through a self-attention recognition strategy, and identifying text information of the feature image through a text recognition algorithm; the target image information is image information of a target image in the feature image;

[0125] determining target image edge information corresponding to the feature image according to the text information of the feature image and the target image information.

[0126] Optionally, the processing module 530 is specifically configured to:

[0127] adding image disturbance noise to the target image information in the target image edge information of each single-frame picture information, and performing blurring processing on the text information in the target image edge information to obtain processed single-frame picture information corresponding to the single-frame picture information;

[0128] combining all the processed single-frame picture information into processed picture information according to the time sequence of the picture information.

[0129] Optionally, the sound elimination module 540 is specifically configured to:

[0130] performing speech recognition processing on the audio information to obtain audio text information corresponding to each audio data in the audio information, and determining whether there is user audio text information in the audio text information;

[0131] In a case where the audio text information is user audio text information, audio interference information is added in audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting a user voiceprint feature of the audio data.

[0132] Optionally, the sound elimination module 540 is specifically configured to:

[0133] According to a preset target information screening strategy, target text information of the audio text information is extracted, and audio data corresponding to the target text information is taken as target audio data.

[0134] The target audio data in the initial processed audio information is subjected to sound elimination processing to obtain processed audio information.

[0135] The above-mentioned various modules in the video processing apparatus can be all or partially realized by software, hardware and combinations thereof. The above-mentioned various modules can be embedded in or independent of a processor in a computer device in a hardware form, or can be stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to the above-mentioned various modules.

[0136] In an embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram thereof can be as shown in Figure 6 The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a video processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or can be a key, trackball or touchpad arranged on the shell of the computer device, or can be an external keyboard, touchpad or mouse, etc.

[0137] Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0138] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:

[0139] Obtaining target video information; the video information comprising a plurality of single-frame picture information and audio information;

[0140] Extracting feature images of each single-frame picture information, and identifying target image edge information in each feature image through a picture edge identification strategy;

[0141] Adding image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining each processed single-frame picture information into processed picture information; the number of processed single-frame picture information is the same as the number of single-frame picture information;

[0142] Adding audio interference information in the audio information, and performing sound elimination processing on target audio data in the audio information that meets the sound elimination processing condition to obtain processed audio information;

[0143] Determining processed video information according to the processed audio information and the processed picture information.

[0144] Optionally, the extracting of the feature images of each single-frame picture information comprises:

[0145] Frame processing the picture information to obtain a plurality of single-frame picture information;

[0146] For each single-frame picture information, the single-frame picture information is divided into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information;

[0147] Extracting initial feature images of each picture information of adjacent single-frame picture information in time sequence, integrating the initial feature images of adjacent picture information according to time sequence, and extracting dynamic interaction features of the integrated initial feature images to obtain feature images of adjacent single-frame picture information.

[0148] Optionally, the identifying of the target image edge information in each feature image through the picture edge identification strategy comprises:

[0149] For each feature image, identifying target image information in the feature image through a self-attention identification strategy, and identifying text information of the feature image through a text recognition algorithm; the target image information is image information of a target image in the feature image;

[0150] According to the text information of the feature image and the target image information, the target image edge information corresponding to the feature image is determined.

[0151] Optionally, image disturbance noise is added in the image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and each processed single-frame picture information is combined into processed picture information, including:

[0152] Image disturbance noise is added in the target image information in the target image edge information of each single-frame picture information, and the text information in the target image edge information is blurred to obtain processed single-frame picture information corresponding to the single-frame picture information.

[0153] All processed single-frame picture information is combined into processed picture information according to the time sequence of the picture information.

[0154] Optionally, the audio interference information is added in the audio information, including:

[0155] The audio information is subjected to speech recognition processing to obtain audio text information corresponding to each audio data in the audio information, and it is judged whether there is user audio text information in the audio text information.

[0156] In the case that there is user audio text information in the audio text information, audio interference information is added in the audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting the user voiceprint feature of the audio data.

[0157] Optionally, the audio information is subjected to speech recognition processing to obtain audio text information corresponding to each audio data in the audio information, and it is judged whether there is user audio text information in the audio text information.

[0158] According to a preset target information screening strategy, target text information of the audio text information is extracted, and audio data corresponding to the target text information is taken as target audio data.

[0159] The target audio data in the initial processed audio information is subjected to speech recognition processing to obtain processed audio information.

[0160] In one embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is executed by a processor to implement the following steps:

[0161] Target video information is obtained; the video information includes a plurality of single-frame picture information and audio information.

[0162] extract a feature image of each single-frame picture information, and identify target image edge information in each feature image through a picture edge identification strategy;

[0163] add image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combine each processed single-frame picture information into processed picture information; the number of the processed single-frame picture information is the same as the number of the single-frame picture information;

[0164] add audio interference information to the audio information, and perform mute processing on target audio data in the audio information that meets a mute processing condition to obtain processed audio information;

[0165] determine processed video information according to the processed audio information and the processed picture information.

[0166] Optionally, the extracting of the feature image of each single-frame picture information comprises:

[0167] frame processing is performed on the picture information to obtain a plurality of single-frame picture information;

[0168] For each single-frame picture information, the single-frame picture information is divided into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information;

[0169] extract initial feature images of picture information of adjacent single-frame picture information in a time sequence, integrate the initial feature images of adjacent picture information according to the time sequence, and extract dynamic interaction features of the integrated initial feature images to obtain feature images of adjacent single-frame picture information.

[0170] Optionally, the identifying of the target image edge information in each feature image through the picture edge identification strategy comprises:

[0171] For each feature image, target image information in the feature image is identified through a self-attention identification strategy, and text information of the feature image is identified through a text recognition algorithm; the target image information is image information of a target image in the feature image;

[0172] determine target image edge information corresponding to the feature image according to the text information of the feature image and the target image information.

[0173] Optionally, the adding of image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and the combining of each processed single-frame picture information into processed picture information, comprises:

[0174] adding image disturbance noise to target image information in target image edge information of each single-frame picture information, and performing blurring processing on text information in the target image edge information, to obtain processed single-frame picture information corresponding to the single-frame picture information;

[0175] combining all the processed single-frame picture information into processed picture information according to a time sequence of the picture information.

[0176] Optionally, the adding audio interference information in the audio information comprises:

[0177] performing speech recognition processing on the audio information to obtain audio text information corresponding to each audio data in the audio information, and determining whether there is user audio text information in the audio text information;

[0178] in a case where there is user audio text information in the audio text information, adding audio interference information in audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting a user voiceprint feature of the audio data.

[0179] Optionally, the muting processing on target audio data in the audio information to obtain processed audio information comprises:

[0180] extracting target text information of the audio text information according to a preset target information screening strategy, and taking audio data corresponding to the target text information as target audio data;

[0181] performing muting processing on target audio data in the initial processed audio information to obtain processed audio information.

[0182] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the following steps:

[0183] obtaining target video information; the video information comprises a plurality of single-frame picture information and audio information;

[0184] extracting a feature image of each single-frame picture information, and identifying target image edge information in each feature image through a picture edge identification strategy;

[0185] adding image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining each processed single-frame picture information into processed picture information; the number of the processed single-frame picture information is the same as the number of the single-frame picture information;

[0186] Adding audio interference information in the audio information, and performing mute processing on target audio data in the audio information satisfying a mute processing condition to obtain processed audio information;

[0187] According to the processed audio information and the processed picture information, determine the processed video information.

[0188] Optionally, the feature image of each single-frame picture information includes:

[0189] Frame processing the picture information to obtain a plurality of single-frame picture information;

[0190] For each single-frame picture information, the single-frame picture information is divided into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information;

[0191] Extract the initial feature image of each picture information of adjacent single-frame picture information in time sequence, and integrate the initial feature image of adjacent picture information according to time sequence, and extract the dynamic interaction feature of the integrated initial feature image to obtain the feature image of adjacent single-frame picture information.

[0192] Optionally, the target image edge information in each feature image is identified by a picture edge identification strategy, including:

[0193] For each feature image, the target image information in the feature image is identified by a self-attention identification strategy, and the text information of the feature image is identified by a text recognition algorithm; the target image information is the image information of the target image in the feature image;

[0194] According to the text information of the feature image and the target image information, determine the target image edge information corresponding to the feature image.

[0195] Optionally, the image disturbance noise is added to the image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and each processed single-frame picture information is combined into processed picture information, including:

[0196] Adding image disturbance noise to the target image information in the target image edge information of each single-frame picture information, and performing blurring processing on the text information in the target image edge information to obtain the processed single-frame picture information corresponding to the single-frame picture information;

[0197] All processed single-frame picture information is combined into processed picture information according to the time sequence of the picture information.

[0198] Optionally, the adding audio interference information in the audio information comprises:

[0199] performing speech recognition processing on the audio information to obtain audio text information corresponding to each audio data in the audio information, and determining whether there is user audio text information in the audio text information;

[0200] in a case where there is user audio text information in the audio text information, adding audio interference information in audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting user voiceprint features of the audio data.

[0201] Optionally, the muting processing on the target audio data in the audio information to obtain the processed audio information comprises:

[0202] extracting target text information of the audio text information according to a preset target information screening strategy, and taking audio data corresponding to the target text information as target audio data;

[0203] performing muting processing on the target audio data in the initial processed audio information to obtain the processed audio information.

[0204] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0205] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0206] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0207] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method of video processing, the method comprising: The method comprises: acquiring target video information; the video information comprises a plurality of single-frame picture information and audio information; extracting feature images of each single-frame picture information, and identifying target image edge information in each feature image through a picture edge identification strategy; adding image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining each processed single-frame picture information into processed picture information; the number of the processed single-frame picture information is the same as the number of the single-frame picture information, and the image disturbance noise is an image adversarial sample that makes the target image information lose a biological feature that can be identified by an identification system; adding audio interference information to the audio information, and performing mute processing on target audio data in the audio information that meets a mute processing condition to obtain processed audio information; determining processed video information according to the processed audio information and the processed picture information; the target image edge information in each feature image is identified through a picture edge identification strategy, comprising: for each feature image, target image information in the feature image is identified through a self-attention identification strategy, and text information of the feature image is identified through a text recognition algorithm; the target image information is image information of a target image in the feature image, the target image is an image related to information that a user needs to protect, and at least includes a face image, an identification card image and an article image; the text information is text information related to personal information that needs to be protected, and at least includes a name, a phone number, an email, an address and an ID card number; determining target image edge information corresponding to the feature image according to the text information of the feature image and the target image information.

2. The method of claim 1, wherein, the feature images of each single-frame picture information are extracted, comprising: frame processing the picture information to obtain a plurality of single-frame picture information; for each single-frame picture information, the single-frame picture information is segmented into a plurality of picture information to obtain a plurality of picture information of each single-frame picture information; extracting initial feature images of each picture information of adjacent single-frame picture information in a time sequence, integrating the initial feature images of adjacent picture information according to the time sequence, and extracting dynamic interaction features of the integrated initial feature images to obtain feature images of adjacent single-frame picture information.

3. The method of claim 1, wherein, adding image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combining each processed single-frame picture information into processed picture information, comprising: adding image disturbance noise to target image information in the target image edge information of each single-frame picture information, and performing blurring processing on text information in the target image edge information to obtain processed single-frame picture information corresponding to the single-frame picture information; combining all processed single-frame picture information into processed picture information according to the time sequence of the picture information.

4. The method of claim 1, wherein, the audio interference information is added to the audio information, comprising: voice recognition processing is performed on the audio information to obtain audio text information corresponding to each audio data in the audio information, and it is determined whether there is user audio text information in the audio text information; In the case where there is user audio text information in the audio text information, audio interference information is added in the audio data corresponding to the user audio text information to obtain initial processed audio information; the interference information is data information affecting the user voiceprint features of the audio data.

5. The method of claim 4, wherein, The method further includes: According to a preset target information screening strategy, target text information of the audio text information is extracted, and audio data corresponding to the target text information is taken as target audio data; The target audio data in the initial processed audio information is subjected to muting processing to obtain the processed audio information.

6. A video processing apparatus, comprising: The device includes: An acquisition module is configured to acquire target video information; the video information includes a plurality of single-frame picture information and audio information; An extraction module is configured to extract feature images of each single-frame picture information, and identify target image edge information in each feature image through a picture edge recognition strategy; A processing module is configured to add image disturbance noise to image information corresponding to each target image edge information to obtain a plurality of processed single-frame picture information, and combine each processed single-frame picture information into processed picture information; the number of the processed single-frame picture information is the same as the number of the single-frame picture information, and the image disturbance noise is an image adversarial sample that causes target image information to lose biological features that can be recognized by a recognition system; An audio muting module is configured to add audio interference information in the audio information, and perform muting processing on target audio data in the audio information that meets a muting processing condition to obtain processed audio information; A determination module is configured to determine processed video information according to the processed audio information and the processed picture information; The extraction module is specifically configured to, for each feature image, identify target image information in the feature image through a self-attention recognition strategy, and identify text information of the feature image through a text recognition algorithm; the target image information is image information of a target image in the feature image, the target image is an image related to information that a user needs to protect, and at least includes a face image, an ID card image, and an article image; the text information is text information related to personal information that needs to be protected, and at least includes a name, a phone number, an email address, an address, and an ID card number; Target image edge information corresponding to the feature image is determined according to the text information of the feature image and the target image information. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method in any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 5.

9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and medium

    CN110555833A

  • Video data silencing method and device

    CN110753262A

  • Method and device for determining special effect video, electronic equipment and storage medium

    CN114630057A