Image desensitization processing method and device, electronic equipment and storage medium

By identifying and filling in clothing materials that blend with adjacent parts in an image, the problem of image desensitization processing destroying integrity in existing technologies is solved, thus improving user experience while ensuring image integrity.

CN114863469BActive Publication Date: 2026-04-24TENCENT TECH (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECH (BEIJING) CO LTD
Filing Date
2021-02-04
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for image desensitization use methods such as mosaic to destroy the integrity of image content, affecting the user's viewing experience, and there is a lack of effective solutions.

Method used

By identifying exposed and sensitive areas in an image and filling them with target clothing material to blend them with adjacent areas, a neural network model is used to determine the style, type, and tolerance of the clothing material, ensuring the integrity of the filled image and the user experience.

Benefits of technology

It achieves accurate masking of sensitive areas while ensuring the integrity of image content, improving the user viewing experience, and the filling process is imperceptible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863469B_ABST
    Figure CN114863469B_ABST
Patent Text Reader

Abstract

The application provides a desensitization processing method and device of an image, electronic equipment and a computer readable storage medium, and relates to cloud technology.The method comprises the following steps: in response to an image display triggering operation, performing identification processing on an image to be displayed; when it is identified that an object in the image has a sensitive part exposed, filling a target clothing material in the sensitive part; wherein the target clothing material is used to be fused with clothing material in an adjacent part, and the adjacent part is a part connected with the sensitive part; and displaying the image after the target clothing material is filled. Through the application, the image can be purified under the premise of ensuring the integrity of the image content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image recognition technology and cloud technology, and more particularly to an image desensitization processing method, apparatus, electronic device and computer-readable storage medium. Background Technology

[0002] Artificial intelligence (AI) is the theory, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. With the research and advancement of AI technology, it has been studied and applied in multiple fields.

[0003] Taking image desensitization as an example, with the development of internet technology, a large number of images (such as photos and videos) containing sensitive content (such as pornography) have appeared on the internet. To purify these images, related technologies typically use materials such as mosaics and preset props to fill in the areas containing sensitive content. However, this method destroys the integrity of the image content and also affects the user's viewing experience. There is currently no effective solution to this problem. Summary of the Invention

[0004] This application provides an image desensitization processing method, apparatus, electronic device, and computer-readable storage medium, which can purify images while ensuring the integrity of image content.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides an image desensitization processing method, including:

[0007] In response to an image display trigger operation, the image to be displayed is identified and processed;

[0008] When an object in the image is identified to have exposed sensitive areas, target clothing material is filled into those sensitive areas;

[0009] The target clothing material is used to blend with clothing materials in adjacent areas, where the adjacent areas are those connected to the sensitive areas.

[0010] The image is displayed after the target clothing material has been filled.

[0011] This application provides an image desensitization processing apparatus, comprising:

[0012] The recognition module is used to recognize and process the image to be displayed in response to the image display trigger operation;

[0013] A filling module is used to fill the sensitive areas with target clothing material when an object in the image is identified as having exposed sensitive areas.

[0014] The target clothing material is used to blend with clothing materials in adjacent areas, where the adjacent areas are those connected to the sensitive areas.

[0015] The display module is used to display the image after the target clothing material has been filled.

[0016] In the above scheme, when the image to be displayed is a video frame in a video, the filling module is further configured to obtain clothing material to be filled in the sensitive part of the reference frame when the similarity between the video frame and the reference frame exceeds a similarity threshold, wherein the playback time of the reference frame is earlier than that of the video frame; and use the clothing material to be filled in the sensitive part of the reference frame as the target clothing material to be filled in the sensitive part of the video frame.

[0017] In the above scheme, when the image to be displayed is a video frame in a video, the filling module is further used to determine the part similarity between the sensitive part in the video frame and the sensitive part in the reference frame; when the part similarity exceeds the part similarity threshold, the clothing material filled in the sensitive part of the reference frame is used as the target clothing material for filling the sensitive part of the video frame; wherein, the playback time of the reference frame is earlier than that of the video frame.

[0018] In the above scheme, the filling module is also used to identify clothing material in the adjacent parts; when it is identified that the adjacent parts are filled with clothing material, the clothing material filled in the adjacent parts is used as the target clothing material for filling in the sensitive parts of the image.

[0019] In the above scheme, the filling module is further used to identify clothing material of the object; when it is identified that the object does not have clothing material, a candidate clothing material belonging to the target clothing style is selected from the clothing material library as the target clothing material for filling the object in the image; wherein, the target clothing material is used to fill the sensitive part and the adjacent part.

[0020] In the above scheme, the filling module is further configured to select candidate clothing materials from the clothing material library that meet the fusion conditions and are suitable for the body shape of the object; wherein, the fusion conditions include at least one of the following: the style of the candidate clothing material matches the dressing style of the object in the image; the type of the candidate clothing material matches the feature information of the object, the feature information including at least one of the object's occupation, gender, and personality; the type of the candidate clothing material matches the filling preference type of the viewing user.

[0021] In the above scheme, the filling module is further used to determine the ratio of the exposed sensitive parts to the area of ​​the object as the exposure ratio of the object; when the viewing user's tolerance for the exposure ratio of the object exceeds the tolerance threshold, or when the exposure ratio of the object exceeds the exposure ratio specified in the content review specifications, it is determined that the operation of selecting candidate clothing materials belonging to the target clothing style from the clothing material library will be performed.

[0022] In the above scheme, the recognition module is further configured to call the first neural network model to perform the following processing: extract feature vectors of multiple candidate regions in the image, map each feature vector to the location and confidence level of the exposed sensitive parts in the candidate region, and determine that the location of the candidate region includes exposed sensitive parts when the confidence level exceeds a confidence level threshold; wherein, the first neural network model is trained based on sample candidate regions and labeled data for the sample candidate regions, and the labeled data includes the location and confidence level of the exposed sensitive parts in the sample candidate regions.

[0023] In the above solution, the filling module is further configured to determine the ratio of the area of ​​the exposed sensitive parts to that of the object as the exposure ratio of the object; determine the tolerance of the viewing user for the exposure ratio of the object; when the tolerance exceeds the tolerance threshold, determine the exposure ratio threshold corresponding to the tolerance threshold, and fill the sensitive parts with the target clothing material so that after the target clothing material is filled into the sensitive parts, the exposure ratio of the object does not exceed the exposure ratio threshold.

[0024] In the above scheme, the filling module is further used to call the second neural network model to perform the following processing: extract the feature vector of the sensitive part, and map the extracted feature vector to the tolerance; wherein, the second neural network model is trained using the historical images viewed by the viewing user and the tolerance for the exposure ratio of the sensitive parts in the historical images as samples.

[0025] In the above scheme, the filling module is further used to collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; and to call a third neural network model to perform the following processing: extracting feature vectors of the behavioral images and mapping the extracted feature vectors to tolerance; wherein the third neural network model is trained using sample behavioral images of the viewing user and tolerances labeled for the sample behavioral images as samples.

[0026] In the above scheme, the filling module is further configured to collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; match the behavioral images with template images of different behavioral types, and take the behavioral type corresponding to the successfully matched template image as the behavioral type of the viewing user; based on the identified behavioral type, query the mapping relationship between different behavioral types and tolerance, and obtain the tolerance corresponding to the identified behavioral type.

[0027] This application provides an electronic device, including:

[0028] Memory is used to store executable instructions for a computer;

[0029] The processor, when executing computer-executable instructions stored in the memory, implements the image desensitization processing method provided in the embodiments of this application.

[0030] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image desensitization processing method provided in this application.

[0031] The embodiments of this application have the following beneficial effects:

[0032] By filling the exposed sensitive areas of objects in an image with target clothing material, the sensitive areas after filling with target clothing material are integrated with the clothing material of adjacent areas in the image as a whole. This achieves accurate masking of exposed sensitive areas in the image while ensuring the integrity of the masked image. Moreover, the image is difficult to perceive as having been processed, thereby improving the user's viewing experience. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the architecture of the image desensitization processing system 100 provided in the embodiments of this application;

[0034] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0035] Figure 3A and Figure 3BThis is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application;

[0036] Figure 4 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application;

[0037] Figure 5 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application;

[0038] Figure 6 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application;

[0039] Figure 7 This is a schematic diagram illustrating an application scenario of the image desensitization processing method provided in the embodiments of this application;

[0040] Figure 8 This is a schematic diagram illustrating an application scenario of the image desensitization processing method provided in the embodiments of this application;

[0041] Figure 9 This is a schematic diagram illustrating an application scenario of the image desensitization processing method provided in the embodiments of this application. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0045] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0046] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0047] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0048] 2) Client: An application running on a terminal that provides various services, such as a picture viewing client, a video client, or a live streaming client.

[0049] 3) Clothing material refers to graphic elements that can be superimposed on an image to give the image a new display effect. Specifically, it is used to fill in the exposed sensitive parts of an object (such as a person or animal), so that when the user views the image, the object appears to be fully clothed.

[0050] The types of clothing materials can include: clothing texture materials, clothing color materials, and outerwear materials, etc.

[0051] 4) Image segmentation, or image cutting, refers to the process of subdividing a digital image into multiple image sub-regions (sets of pixels) (also known as superpixels). A superpixel is a small region composed of a series of adjacent pixels with similar characteristics such as color, brightness, and texture.

[0052] The relevant technologies only handle sensitive content in images by blurring it out (for example, using mosaics or special effects to fill in sensitive content in the image). This will cause the filled area in the image to not blend with the adjacent area, making it obvious to the user and thus affecting the user's viewing experience.

[0053] To address the aforementioned technical problems, embodiments of this application provide an image desensitization processing method that can merge the filled area and adjacent areas in an image to ensure the integrity of the content in the image, thereby achieving the effect of purifying the image without being noticed.

[0054] The following describes an exemplary application of the image desensitization processing method provided in the embodiments of this application. The image desensitization processing method provided in the embodiments of this application can be implemented by various electronic devices, such as smartphones, tablets, vehicle terminals, smart wearable devices, laptops, desktop computers, smart TVs, and other types of user terminals (hereinafter referred to as terminals). The electronic device can also be a server.

[0055] Next, taking an electronic device as an example, an exemplary application system architecture for implementing the image desensitization processing method provided in the embodiments of this application will be described. See [link to documentation]. Figure 1 , Figure 1 This is a schematic diagram of the architecture of the image desensitization processing system 100 provided in this application embodiment. The image desensitization processing system 100 includes a server 200, a network 300, and a terminal 400, which will be described separately.

[0056] Server 200 is the backend server for client 410, used to respond to client 410's image acquisition request and send the image to be displayed to client 410.

[0057] Network 300 is used as a medium for communication between server 200 and terminal 400, and can be a wide area network, a local area network, or a combination of both.

[0058] Terminal 400 is used to run client 410, which is a client with image desensitization function, such as an image viewing client, video client, live streaming client, browser client, instant messaging client, educational client, etc. Client 410 is used to respond to the image display trigger operation, obtain the image to be displayed sent by server 200, and perform recognition processing on the image to be displayed; it is also used to fill the sensitive parts of the image with target clothing material when the object in the image is identified as having exposed sensitive parts, and display the image after filling with target clothing material on the human-computer interaction interface.

[0059] In some embodiments, the terminal 400 implements the image desensitization processing method provided in this application embodiment by running a computer program. For example, the computer program may be a native program or software module in an operating system; it may be a native application (APP), i.e., a program that needs to be installed in the operating system to run, such as a local browser; it may also be a mini-program, i.e., a program that only needs to be downloaded into the browser environment to run; or it may be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program may be any form of application, module, or plugin.

[0060] The embodiments of this application can be implemented with the help of cloud technology, which refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.

[0061] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, allowing for on-demand use with flexibility and convenience. Cloud computing technology will become a crucial support. The backend services of cloud computing systems require substantial computing and storage resources.

[0062] As an example, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal 400 and server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0063] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. In this embodiment, the electronic device is used as an example. Figure 1 The following explanation will be based on terminal 400. Figure 2 The terminal 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0064] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0065] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0066] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0067] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0068] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0069] The operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.

[0070] The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB).

[0071] Presentation module 453 enables the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with user interface 430 (e.g., a display screen, a speaker, etc.).

[0072] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0073] In some embodiments, the image desensitization processing apparatus provided in this application can be implemented in software. Figure 2 A desensitization processing apparatus 455 for images stored in memory 450 is shown. This apparatus can be software in the form of programs and plug-ins, and includes the following software modules: an identification module 4551, a filling module 4552, and a display module 4553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0074] It should be noted that, based on the... Figure 2 In this understanding, electronic devices can correspond to servers with similar functional structures, for example, Figure 1 Server 200 in the middle. When Figure 2The electronic device shown is Figure 1 When the server is 200, Figure 2 The user interface 430, presentation module 453 and input processing module 454 shown can be omitted.

[0075] The image desensitization processing method provided in this application embodiment can be derived from... Figure 1 Terminal 400 or server 200 can be executed independently, or it can be... Figure 1 Terminal 400 and server 200 work together to execute.

[0076] Below, by Figure 1 The image desensitization processing method provided in this application embodiment is illustrated using terminal 400 in the example. See also... Figure 3A , Figure 3A This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application, which will be combined with Figure 3A The steps shown are explained.

[0077] It should be noted that, Figure 3A The method shown can be executed by various forms of computer programs running on terminal 400, and is not limited to the client 410 described above, such as the operating system 451, software modules and scripts mentioned above. Therefore, the client example in the following text should not be regarded as a limitation on the embodiments of this application.

[0078] In step S101, in response to the image display trigger operation, the image to be displayed is subjected to recognition processing.

[0079] In some embodiments, the image to be displayed may be a picture, a video frame in a video (e.g., a pre-recorded video), or a real-time video frame (e.g., a video frame during a live broadcast).

[0080] Taking the image to be displayed as an example, in response to the image viewing operation, the image to be displayed is identified to determine whether the object in the image (such as a person or an animal) has exposed sensitive parts (such as the chest or buttocks).

[0081] Taking a video frame from a video as an example, the video frame can be a video frame that has been decoded from the video but has not yet been displayed. In response to the video viewing operation, the video is decoded to obtain multiple video frames. For each video frame, identification processing is performed to determine whether the object in the video frame has exposed sensitive parts.

[0082] Taking the image to be displayed as a real-time video frame during a live broadcast as an example, in response to the live broadcast viewing operation, the live broadcast data is acquired and decoded to obtain multiple real-time video frames. For each real-time video frame, identification processing is performed to determine whether the object in the real-time video frame has exposed sensitive parts.

[0083] In some embodiments, the client can invoke the corresponding service of the terminal (e.g., image recognition service) to complete the image recognition process through the terminal. Alternatively, the client can invoke the corresponding service of the server (e.g., image recognition service) to complete the image recognition process through the server.

[0084] As an example, when the client calls the corresponding service of the server (e.g., image recognition service) to complete the image recognition process, the alternative step of step S101 is: the client responds to the image display trigger operation and sends an image recognition instruction to the server so that the server can perform recognition processing on the image to be displayed.

[0085] The following explanation uses the example of a client calling a corresponding service (e.g., an image recognition service) on a terminal to complete image recognition. It should be noted that the process of a client calling a corresponding service (e.g., an image recognition service) on a server to complete image recognition is similar and will not be repeated here.

[0086] In some embodiments, the first neural network model is invoked to perform the following processing: extract feature vectors of multiple candidate regions in the image, map each feature vector to the location of the candidate region including the exposed sensitive parts and the confidence level, and determine that the candidate region includes the exposed sensitive parts when the confidence level exceeds the confidence level threshold.

[0087] As an example, the training samples for the first neural network model can be candidate regions of samples and labeled data for the candidate regions of samples, such as the location of exposed sensitive areas and the confidence level of exposed sensitive areas.

[0088] As an example, the first neural network model can be of various types, such as a Convolutional Neural Network (CNN) model, a Recurrent Neural Network (RNN) model, and a multilayer feedforward neural network model. The aforementioned neural network model can be trained using a supervised approach, where the loss function used to train the first neural network model represents the difference between the predicted value and the actual labeled data. This loss function can be a 0-1 loss function, a perceptual loss function, or a cross-entropy loss function, etc.

[0089] This application embodiment uses a neural network model to determine the location of sensitive parts in an image, which can improve the accuracy of determining the location of sensitive parts, thereby improving the efficiency of subsequent filling of exposed sensitive parts.

[0090] In step S102, when it is identified that an object in the image has exposed sensitive parts, the target clothing material is filled into the sensitive parts.

[0091] Here, the target clothing material is used to blend with clothing materials in adjacent areas, which are the areas connected to the sensitive areas.

[0092] In some embodiments, the client can invoke the corresponding service of the terminal (e.g., material filling service) to complete the material filling process through the terminal. Alternatively, the client can invoke the corresponding service of the server (e.g., material filling service) to complete the material filling process through the server.

[0093] As an example, when the client calls the corresponding service of the server (e.g., material filling service) to complete the material filling process, the alternative step of step S102 is: when the object in the image is identified to have exposed sensitive parts, the server fills the sensitive parts with target clothing material; and sends the image after filling with target clothing material to the client.

[0094] The following explanation uses the example of a client calling a corresponding service on the terminal (e.g., a content filling service) to complete content filling. It should be noted that the process of a client calling a corresponding service on the server (e.g., a content filling service) to complete content filling is similar and will not be repeated here.

[0095] In some embodiments, the ratio of the exposed sensitive area to the area of ​​the object is determined as the exposure ratio of the object; the tolerance of the viewing user for the exposure ratio of the object is determined; when the tolerance exceeds the tolerance threshold, the exposure ratio threshold corresponding to the tolerance threshold is determined, and the target clothing material is filled into the sensitive area so that after the target clothing material is filled into the sensitive area, the exposure ratio of the object does not exceed the exposure ratio threshold.

[0096] As an example, viewers can be either logged-in users or viewers viewing images. Each viewer has a different tolerance level for different proportions of nudity, and there is a positive correlation between the proportion of nudity and tolerance. For example, viewer A has a tolerance of 0.3 for a 30% nudity proportion and a tolerance of 0.4 for a 40% nudity proportion.

[0097] As an example, the tolerance threshold and the exposure ratio threshold can be default values ​​or values ​​set by the user, client, or server.

[0098] As an example, in an image, areas containing sensitive parts are filled with target clothing material; the area of ​​the area to be filled is negatively correlated with the tolerance level. Thus, the higher the viewer's tolerance level, the smaller the area to be filled, allowing for personalized image filling. This not only saves filling resources but also improves the user's viewing experience.

[0099] For example, Figure 7 In this context, user A has a lower tolerance for the original image than user B. Therefore, the area of ​​the region to be filled 701 set for user A is larger than the area of ​​the region to be filled 701 set for user B.

[0100] The following example illustrates how to determine the tolerance level of a viewing user for the proportion of nudity an object should have.

[0101] As a first example, the second neural network model is invoked to perform the following processing: extract feature vectors of sensitive areas and map the extracted feature vectors to tolerance.

[0102] For example, the second neural network model was trained using samples of historical images viewed by users and their tolerance for the proportion of nudity in sensitive areas of those historical images.

[0103] For example, the second neural network model can include various types, such as convolutional neural network models, recurrent neural network models, and multilayer feedforward neural network models. These neural network models can be trained using supervised methods. The loss function used to train the second neural network model represents the difference between the predicted value and the actual labeled data; the loss function can be a 0-1 loss function, a perceptual loss function, or a cross-entropy loss function, etc.

[0104] This application embodiment uses a neural network model to determine the viewer's tolerance for sensitive parts in an image, which can improve the accuracy of the determined tolerance and personalize the image filling according to the tolerance. This not only saves filling resources but also improves the efficiency of image filling.

[0105] As a second example, a camera device (e.g., a terminal's camera) is invoked to capture images of the user's behavior, wherein the behavior images include at least one of facial expressions and body movements; a third neural network model is invoked to perform the following processing: extracting feature vectors from the behavior images and mapping the extracted feature vectors to tolerance.

[0106] For example, the third neural network model is trained using images of user behavior samples and the tolerance for annotations on those images as samples.

[0107] For example, the third neural network model can include various types, such as convolutional neural network models, recurrent neural network models, and multilayer feedforward neural network models. These neural network models can be trained using supervised methods. The loss function used to train the third neural network model represents the difference between the predicted value and the actual labeled data; the loss function can be a 0-1 loss function, a perceptual loss function, or a cross-entropy loss function, etc.

[0108] For example, behavioral images may include eye behavior, head behavior, voice behavior, or gesture behavior.

[0109] This application embodiment determines the tolerance level of the corresponding viewing user's behavior image through a neural network model, which can improve the accuracy of the determined tolerance level, and fill the image in a personalized way according to the tolerance level, which can not only save filling resources, but also improve the efficiency of image filling.

[0110] As a third example, a camera device (e.g., a terminal's camera) is invoked to capture images of the viewing user's behavior, wherein the behavior images include at least one of facial expressions and body movements; the behavior images are matched with template images of different behavior types, and the behavior type corresponding to the successfully matched template images is taken as the viewing user's behavior type; based on the identified behavior type, the mapping relationship between different behavior types and tolerance is queried to obtain the tolerance corresponding to the identified behavior type.

[0111] For example, behavioral images may include eye behavior, head behavior, voice behavior, or gesture behavior.

[0112] For example, the tolerance for eye behaviors such as squinting, body behaviors such as covering one's vision with one's hands or objects, voice behaviors that indicate the emotion type as disgust after speech recognition of voice information (such as "disgusted" or "very nauseous"), or head behaviors such as frowning and downturned corners of the mouth is higher than the tolerance for flat, unwavering facial expressions or body movements.

[0113] This application embodiment determines the tolerance of the corresponding viewing user's behavior image through a mapping relationship, which can improve the speed of determining the tolerance and personalize the image filling according to the tolerance. This not only saves filling resources but also improves the efficiency of image filling.

[0114] In some embodiments, when the image to be displayed is a video frame that has been decoded from the video but has not yet been displayed, before filling the sensitive area with the target clothing material, the clothing material to be filled in the sensitive area of ​​the reference frame can be obtained when the similarity between the video frame and the reference frame exceeds a similarity threshold; the clothing material to be filled in the sensitive area of ​​the reference frame is used as the target clothing material to be filled in the sensitive area of ​​the video frame.

[0115] As an example, the playback time of the reference frame is earlier than that of the video frame. For example, the reference frame can be the video frame preceding the video frame, or it can be the video frame preceding the previous N (N is an integer greater than 1) video frames. This application embodiment does not impose any restrictions on this.

[0116] As an example, the similarity threshold can be a default value, a value set by the user, client, or server, or it can be determined based on the similarity between all adjacent video frames. For example, the average similarity between all adjacent video frames can be used as the similarity threshold.

[0117] As an example, determining the similarity between a video frame and a reference frame can be achieved by calling a fourth neural network model to perform the following process: extracting feature vectors from both the video frame and the reference frame, and mapping these feature vectors to similarity scores; wherein the fourth neural network model is trained using sample images (e.g., multiple paired images) and similarity scores labeled for the sample images. This improves the accuracy of similarity judgment.

[0118] For example, the fourth neural network model can include various types, such as convolutional neural network models, recurrent neural network models, and multilayer feedforward neural network models. These neural network models can be trained using supervised methods. The loss function used to train the fourth neural network model represents the difference between the predicted value and the actual labeled data; the loss function can be a 0-1 loss function, a perceptual loss function, or a cross-entropy loss function, etc.

[0119] As an example, determining the similarity between a video frame and a reference frame can also involve determining the histograms of the video frame and the reference frame, calculating the normalized correlation coefficient (e.g., Bach distance or histogram intersection distance) between the histograms of the video frame and the reference frame, and determining the similarity based on the normalized correlation coefficient. The normalized correlation coefficient and the similarity are positively correlated. This approach, compared to neural network models, can improve the speed of determining part similarity.

[0120] As an example, clothing elements filling sensitive areas in multiple reference frames are cached. The similarity between the video frame and each reference frame is determined, and the reference frame with the highest similarity is selected. The clothing elements filling sensitive areas in the most similar reference frame are then used as the target clothing elements for filling sensitive areas in the video frame. This saves resources on finding target elements, thereby improving the efficiency of area filling.

[0121] As an example, a video might contain exposed sensitive areas that appear repeatedly in multiple video frames, but the location of these areas differs in each frame. In this case, different sensitive areas and their corresponding target clothing materials from multiple reference frames can be cached. When the similarity between a video frame and a reference frame does not exceed a similarity threshold, the similarity between the sensitive areas included in the video frame and each cached sensitive area is determined. The target clothing material corresponding to the sensitive area with the highest similarity exceeding the similarity threshold is used as the target clothing material for filling the sensitive areas in the video frame. This saves resources searching for target materials, thereby improving the efficiency of area filling.

[0122] For example, the part similarity threshold can be a default value, a value set by the user, client, or server, or it can be determined based on the part similarity between the sensitive parts included in the video frame and all cached sensitive parts. For example, the average part similarity between the sensitive parts included in the video frame and all cached sensitive parts can be used as the part similarity threshold.

[0123] As an example, different sensitive areas and corresponding target clothing materials from multiple reference frames are cached. When the similarity between the video frame and the reference frames does not exceed a similarity threshold, the video frame is divided into multiple candidate regions. The similarity between the region containing the sensitive area in the reference frame and each candidate region is determined. The clothing material corresponding to the region containing the sensitive area in the reference frame with the highest similarity exceeding the similarity threshold is used as the target clothing material for filling the sensitive areas in the video frame. In this way, target material search resources can be saved, thereby improving region filling efficiency.

[0124] For example, the region similarity threshold can be a default value, a value set by the user, client, or server, or it can be determined based on the similarity between all candidate regions and regions containing sensitive parts in the reference frame. For instance, the average similarity between all candidate regions and sensitive parts in the reference frame can be used as the region similarity threshold.

[0125] In other embodiments, when the image to be displayed is a video frame that has been decoded from the video but has not yet been displayed, before filling the sensitive parts with target clothing material, the similarity between the sensitive parts in the video frame and the sensitive parts in the reference frame can be determined; when the similarity exceeds the similarity threshold, the clothing material filled in the sensitive parts of the reference frame is used as the target clothing material to be filled in the sensitive parts of the video frame; wherein, the playback time of the reference frame is earlier than that of the video frame.

[0126] As an example, the target clothing material is cached for sensitive areas in a video frame; when the sensitive area reappears in a subsequent video frame, the target clothing material is filled into the sensitive area.

[0127] As an example, the playback time of the reference frame is earlier than that of the video frame. For example, the reference frame can be the video frame preceding the video frame, or it can be the video frame preceding the previous N (N is an integer greater than 1) video frames. This application embodiment does not impose any restrictions on this.

[0128] As an example, determining the similarity between sensitive parts in a video frame and sensitive parts in a reference frame can be achieved by calling a fifth neural network model to perform the following processing: extracting feature vectors representing sensitive parts from both the video frame and the reference frame, and mapping these extracted feature vectors to part similarity scores; wherein, the fifth neural network model is trained using sample images (e.g., multiple paired images) and part similarity scores labeled for the sample images. This improves the accuracy of part similarity determination.

[0129] For example, the fifth neural network model can include various types, such as convolutional neural network models, recurrent neural network models, and multilayer feedforward neural network models. These neural network models can be trained using supervised methods. The loss function used to train the fifth neural network model represents the difference between the predicted value and the actual labeled data; the loss function can be a 0-1 loss function, a perceptual loss function, or a cross-entropy loss function, etc.

[0130] As an example, determining the similarity between sensitive parts in a video frame and sensitive parts in a reference frame can also involve identifying histograms of regions containing sensitive parts in both the video and reference frames, calculating a normalized correlation coefficient (e.g., Bach distance or histogram intersection distance) between the histograms of regions containing sensitive parts in the video and reference frames, and determining the part similarity based on the normalized correlation coefficient. The normalized correlation coefficient and part similarity are positively correlated. This approach, compared to neural network models, can improve the speed of part similarity determination.

[0131] In step S103, an image of the target clothing material after filling is displayed.

[0132] In some embodiments, see Figure 7In the original image, the figure's chest is partially covered. Therefore, the material used for the clothing on the chest is used to fill the area above and / or below the chest, so that the filled area blends with the area of ​​the original clothing on the chest. Compared with the blurring in related technologies (e.g., using mosaic to fill the image), the embodiments of this application make it visually impossible for users to discover that the figure was originally naked, thereby improving the integrity of the desensitized image and enhancing the user's viewing experience.

[0133] In some embodiments, the displayed image can be a picture, a video frame from a video (e.g., a pre-recorded video), or a real-time video frame (e.g., a video frame during a live broadcast). The following explanation uses the example of a real-time video frame during a live broadcast to illustrate this.

[0134] As an example, when the anchor has exposed sensitive parts in the live video frame during a live broadcast, the sensitive parts are filled with target clothing material, and the live video frame after the target clothing material is filled is displayed.

[0135] For example, when the streamer has exposed sensitive parts in the real-time video frame during a live broadcast, the ratio of the exposed sensitive parts to the streamer's area is determined as the streamer's exposure ratio; the tolerance of the live broadcast audience for the streamer's exposure ratio is determined; when the tolerance exceeds the tolerance threshold, the exposure ratio threshold corresponding to the tolerance threshold is determined, and target clothing material is filled into the sensitive parts so that the streamer's exposure ratio does not exceed the exposure ratio threshold after the target clothing material is filled into the sensitive parts.

[0136] Here, the tolerance level of the live stream audience for the proportion of nudity shown by the streamer is determined in the same way as in the above embodiments, and will not be repeated here.

[0137] For example, different viewers may have different tolerance levels. As a result, the area of ​​the target clothing material filled in for the streamer in the live broadcast screen will be different for different viewers. In other words, the live broadcast screen can be personalized, thereby improving the user's viewing experience.

[0138] In some embodiments, see Figure 3B , Figure 3B This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application, based on Figure 3A Step S104 may be included before step S102.

[0139] In step S104, the target clothing material is determined.

[0140] In some embodiments, clothing material is identified in adjacent areas; when clothing material is identified in adjacent areas, the clothing material in adjacent areas is used as the target clothing material for filling in sensitive areas of the image.

[0141] As an example, Figure 8 In the process, clothing material in adjacent areas 801 is identified. When clothing material is filled in adjacent areas 801, it is used as the target clothing material to fill sensitive areas 802. In this way, the user cannot visually detect that the object originally had exposed sensitive areas, thereby improving the integrity of the desensitized image and enhancing the user's viewing experience.

[0142] In other embodiments, clothing material recognition is performed on the object; when it is identified that the object does not have clothing material, or the type of clothing material that the object has is underwear, candidate clothing material set by the user, client or server is selected from the clothing material library as target clothing material to fill the object in the image, wherein the target clothing material is used to fill sensitive parts and adjacent parts.

[0143] For example, when the viewers are minors, parents can pre-set the candidate clothing materials to be cartoon clothing materials, which can not only increase the fun of watching the image, but also personalize the image.

[0144] In some other embodiments, clothing material recognition is performed on the object; when it is identified that the object does not have clothing material, or the type of clothing material that the object has is underwear, candidate clothing materials (such as clothing texture material, clothing color material and outer garment material) belonging to the target clothing style are selected from the clothing material library as target clothing materials for filling the object in the image; wherein, the target clothing material is used to fill sensitive parts and adjacent parts.

[0145] As an example, candidate clothing materials that meet the fusion conditions and are suitable for the body shape of the object are selected from the clothing material library; wherein, the fusion conditions include at least one of the following: the style of the candidate clothing material matches the dressing style of the object in the image; the type of the candidate clothing material matches the feature information of the object, and the feature information includes at least one of the object's occupation, gender, and personality; the type of the candidate clothing material matches the fill preference type of the viewing user.

[0146] For example, the style of an object in an image can be set by the user, the client, or the server. For instance, a user might set a specific style for each character in a movie.

[0147] For example, the style of an object's appearance in an image can also be determined by style recognition of the object in the image.

[0148] For example, by calling the sixth neural network model to perform the following processing: extracting feature vectors representing the dressing style of an object in the image, and mapping the extracted feature vectors to dressing styles; wherein, the sixth neural network model is trained using sample images and the dressing styles labeled for objects in the sample images as samples. In this way, the accuracy of object dressing style judgment can be improved.

[0149] For example, matching an object's appearance in an image with different template styles, and using the successfully matched template styles as the object's appearance style in the image. This can improve the speed of object style determination.

[0150] For example, Figure 9 In the clothing material library 903, there are candidate clothing materials for skirts 901 and candidate clothing materials for coats 902. Since the style of the person's dress in the original image is ladylike, candidate clothing material for skirts 901 is selected to cover the sensitive areas of the person in the original image.

[0151] For example, the type of candidate clothing material can be determined based on the object's occupation, gender, and personality. For instance, when the person in the original image is a white-collar worker, a professional-style clothing material can be selected from the clothing material library as the target clothing material, rather than a street hip-hop style clothing material.

[0152] For example, the fill preference type for a viewing user can be a fill preference type set by the user, the client, or the server. For instance, a user might set a corresponding fill preference type for each character in a movie.

[0153] As an example, before selecting candidate clothing materials belonging to the target clothing style from the clothing material library, the ratio of the exposed sensitive parts to the area of ​​the object can be determined as the object's exposure ratio. When the viewing user's tolerance for the object's exposure ratio exceeds the tolerance threshold, or when the object's exposure ratio exceeds the exposure ratio stipulated in the content review guidelines, it is determined that the operation of selecting candidate clothing materials belonging to the target clothing style from the clothing material library will be performed.

[0154] For example, determining the viewer's tolerance for the proportion of exposed objects is the same as in the above embodiments and will not be repeated here. If the viewer's tolerance for the proportion of exposed objects exceeds a tolerance threshold, it indicates that the degree of exposure of the image to be displayed exceeds the viewer's tolerance range. In this case, sensitive exposed areas need to be filled in so that the viewer does not reject the subsequently presented image.

[0155] For example, if the exposed area of ​​an object exceeds the exposure ratio specified in the content review guidelines, it means that the exposed area of ​​the object in the image to be displayed is too large, causing the image to fail the content review. In this case, the exposed sensitive areas need to be filled in in order to pass the content review.

[0156] The embodiments of this application determine whether to fill the image based on the exposed ratio of the object, thereby achieving diversity in image filling while saving filling resources.

[0157] Below, by Figure 1 The image desensitization processing method provided in this application embodiment is illustrated by example, in which the terminal 400 and server 200 collaboratively implement the method. See also Figure 4 , Figure 4 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application, which will be combined with Figure 4 The steps shown are explained.

[0158] In step S401, the terminal responds to the image display trigger operation by sending an image acquisition request to the server.

[0159] In step S402, the server acquires the corresponding image based on the image acquisition request and performs recognition processing on the image.

[0160] In step S403, when an object in the image is identified to have exposed sensitive parts, the server fills the sensitive parts with target clothing material.

[0161] In step S404, the server sends the image after filling the target clothing material to the terminal.

[0162] In step S405, the terminal displays an image after the target clothing material has been filled.

[0163] It should be noted that the specific implementation methods of steps S401 to S405 are similar to those of steps S101 to S103 above, and will not be repeated here.

[0164] In this embodiment, the server has stronger computing power and faster processing speed than the terminal. By completing the image filling process through the server, the speed at which the terminal displays the filled image can be improved, and the computing resources of the terminal can be reduced.

[0165] The following example illustrates the image desensitization processing method provided in this application, taking a video frame from a video as the image to be displayed and the object as a person.

[0166] This application embodiment identifies nude scenes in a video, determines the clothing material of the people in the video, and automatically fills the corresponding clothing material into the exposed sensitive areas, achieving seamless occlusion of exposed sensitive areas and thus improving the user's viewing experience. Specifically, firstly, the video stream is detected to identify whether there are exposed sensitive areas in a single frame of the video (i.e., the image mentioned above); then, the clothing of the people in the frame with exposed sensitive areas is identified, and a repair tool is used to fill the exposed sensitive areas in the frame with the corresponding clothing material; finally, the processed single frame is merged into the video stream and re-encoded.

[0167] See Figure 5 , Figure 5 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application. Figure 5 The image desensitization processing method shown can be executed by the terminal or the server alone, or by the terminal and the server working together. The following description uses the image desensitization processing method provided in the embodiments of this application executed by the server as an example.

[0168] In step S501, the server performs single-frame decoding on the source video stream, performs nudity detection on the single-frame image, and determines whether there are any exposed sensitive parts in the single-frame image.

[0169] In step S502, when a single frame contains exposed sensitive parts, the server performs clothing detection near the exposed area to identify the clothing material of the person.

[0170] In step S503, the server fills the exposed area with clothing material and re-encodes the filled single frame into the video stream.

[0171] The specific implementation of the image desensitization processing method provided in the embodiments of this application will be described in detail below.

[0172] (a) Detection of exposed sensitive areas

[0173] In some embodiments, the video is saved frame by frame as an image, for example, by using video tools such as OpenCV or ffmpeg to save the video frame by frame as an image.

[0174] As an example, when the video format is test.mp4, the format to save the video frame by frame as an image using the video tool ffmpeg is ffmpeg -i test.mp4 frames / frames_%05d.jpg.

[0175] As an example, the code for saving video frame by frame as images using the video tool OpenCV could be:

[0176] from __future__ import absolute __import

[0177] from __future__ import division

[0178] from __future__ importprint _function

[0179] import cv2

[0180] cap = cv2.VideoCapture('test.mp4'); Input the video to be processed.

[0181] Success, image cap. read(); reads the video to be processed.

[0182] count = 0

[0183] success = True

[0184] while success:

[0185] # Save frame as JPEG file; saves the image as a JPEG file.

[0186] cv2.imwrite("frames / framegd.jpg" count, image); saves the image to the frames folder.

[0187] success, image = cap. read()

[0188] print (' Read a new frame: ', count)

[0189] count += 1

[0190] In this way, the video can be converted into images frame by frame using the above code, and then detection can be performed on individual frames.

[0191] In some embodiments, images with a lot of exposed skin or skin-colored areas are prone to false detection in related technologies. To reduce the false detection rate for such images, this application extracts skin color features, histogram of oriented gradients (HOG) features that characterize the appearance and shape of local objects, spatial distribution features, and Haar-like features that describe the gray-level distribution of the region. Using the Adaboost learning algorithm, a classifier for sensitive parts of the human body is trained. The classifier is used to filter sensitive images. This application can accurately detect sensitive parts and effectively reduce the false detection rate of sensitive images.

[0192] (ii) Determine clothing materials and area filling

[0193] In some embodiments, after the detection of exposed sensitive areas is completed, an image containing the exposed sensitive areas is obtained; the clothing around the sensitive areas of the person is identified, and an edge similarity filling method is used to fill the exposed sensitive areas with similar clothing material to complete the coverage of the clothing material of the exposed sensitive areas.

[0194] As an example, see Figure 6 , Figure 6 This is a schematic flowchart of the image desensitization processing method provided in the embodiments of this application, based on Figure 5 Step S503 may further include: the server performs superpixelation on the image, uses image segmentation algorithm and approximate Gaussian mixture clustering to segment clothing pixels, and then uses clothing pixels to cover exposed sensitive areas through edge similarity, thereby achieving clothing coverage of exposed sensitive areas.

[0195] As an example, all the processed single-frame images are recombined into a video stream. For instance, the process of combining single-frame images into a video stream using the video tool OpenCV could be as follows: determine the frame rate, frame count, and resolution of the video to be combined; create a video file; load multiple images sequentially according to the playback order; create images of the same size as the video file format; set the loaded images to have the same size as the video format; and save the images as a video stream format.

[0196] This application embodiment can cover up the exposed sensitive parts in movies and TV dramas with clothing. Compared with the blurring in related technologies (e.g., using mosaic to fill the image), this application embodiment is more easily accepted by the user's senses, thereby improving the user's video viewing experience.

[0197] The following is combined Figure 2 The implementation of the image desensitization processing apparatus provided in this application embodiment is an exemplary structure of a software module.

[0198] In some embodiments, such as Figure 2 As shown, the software module in the image desensitization processing device 455 storing the image in the memory 450 may include:

[0199] The recognition module 4551 is used to perform recognition processing on the image to be displayed in response to the image display trigger operation;

[0200] The filling module 4552 is used to fill the sensitive parts of an object in an image with target clothing material when the object is identified to have exposed sensitive parts.

[0201] Among them, the target clothing material is used to blend with the clothing material in the adjacent parts, which are the parts connected to the sensitive parts;

[0202] Display module 4553 is used to display the image after filling the target clothing material.

[0203] In the above scheme, when the image to be displayed is a video frame in a video, the filling module 4552 is further used to obtain clothing material to be filled in the sensitive part of the reference frame when the similarity between the video frame and the reference frame exceeds the similarity threshold, wherein the playback time of the reference frame is earlier than that of the video frame; and to use the clothing material to be filled in the sensitive part of the reference frame as the target clothing material to be filled in the sensitive part of the video frame.

[0204] In the above scheme, when the image to be displayed is a video frame in a video, the filling module 4552 is also used to determine the part similarity between the sensitive part in the video frame and the sensitive part in the reference frame; when the part similarity exceeds the part similarity threshold, the clothing material filled in the sensitive part of the reference frame is used as the target clothing material to be filled in the sensitive part of the video frame; wherein, the playback time of the reference frame is earlier than that of the video frame.

[0205] In the above scheme, the filling module 4552 is also used to identify clothing material in adjacent parts; when it is identified that clothing material is filled in adjacent parts, the clothing material filled in adjacent parts is used as the target clothing material for filling in sensitive parts of the image.

[0206] In the above scheme, the filling module 4552 is also used to identify clothing material of the object; when the object is identified as not having clothing material, candidate clothing material belonging to the target clothing style is selected from the clothing material library as the target clothing material for filling the object in the image; wherein, the target clothing material is used to fill sensitive parts and adjacent parts.

[0207] In the above scheme, the filling module 4552 is further used to select candidate clothing materials from the clothing material library that meet the fusion conditions and are suitable for the body shape of the object; wherein, the fusion conditions include at least one of the following: the style of the candidate clothing material matches the dressing style of the object in the image; the type of the candidate clothing material matches the feature information of the object, and the feature information includes at least one of the object's occupation, gender, and personality; the type of the candidate clothing material matches the filling preference type of the viewing user.

[0208] In the above scheme, the filling module 4552 is also used to determine the ratio of the exposed sensitive parts to the area of ​​the object as the exposure ratio of the object; when the viewing user's tolerance for the exposure ratio of the object exceeds the tolerance threshold, or when the exposure ratio of the object exceeds the exposure ratio specified in the content review specifications, it is determined that the operation of selecting candidate clothing materials belonging to the target clothing style from the clothing material library will be performed.

[0209] In the above scheme, the recognition module 4551 is also used to call the first neural network model to perform the following processing: extract feature vectors of multiple candidate regions in the image, map each feature vector to the location of the candidate region including the exposed sensitive parts and the confidence level, and determine that the candidate region includes the exposed sensitive parts when the confidence level exceeds the confidence level threshold; wherein, the first neural network model is trained based on the sample candidate regions and the labeled data for the sample candidate regions, and the labeled data includes the location of the exposed sensitive parts in the sample candidate regions and the confidence level.

[0210] In the above scheme, the filling module 4552 is also used to determine the area ratio of the exposed sensitive parts to the object as the object's exposure ratio; determine the viewing user's tolerance for the object's exposure ratio; when the tolerance exceeds the tolerance threshold, determine the exposure ratio threshold corresponding to the tolerance threshold, and fill the sensitive parts with target clothing material so that after filling the sensitive parts with target clothing material, the object's exposure ratio does not exceed the exposure ratio threshold.

[0211] In the above scheme, the filling module 4552 is also used to call the second neural network model to perform the following processing: extract the feature vector of the sensitive part and map the extracted feature vector to the tolerance; wherein, the second neural network model is trained using the historical images viewed by the user and the tolerance for the exposure ratio of the sensitive parts in the historical images as samples.

[0212] In the above scheme, the filling module 4552 is also used to collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; and call the third neural network model to perform the following processing: extract the feature vector of the behavioral image, and map the extracted feature vector to tolerance; wherein the third neural network model is trained using sample behavioral images of the viewing user and the tolerance labeled for the sample behavioral images as samples.

[0213] In the above scheme, the filling module 4552 is also used to collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; match the behavioral images with template images of different behavioral types, and take the behavioral type corresponding to the successfully matched template image as the behavioral type of the viewing user; according to the identified behavioral type, query the mapping relationship between different behavioral types and tolerance, and obtain the tolerance corresponding to the identified behavioral type.

[0214] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image desensitization processing method described in this application.

[0215] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they cause the processor to perform the image desensitization processing method provided in this application. For example... Figure 3A , Figure 3B , Figure 4 , Figure 5 and Figure 6 The image desensitization processing method shown is used in computers including various computing devices such as smart terminals and servers.

[0216] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0217] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0218] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hypertext Markup Language document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0219] As an example, computer-executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0220] In summary, the embodiments of this application have the following beneficial effects:

[0221] (1) By filling the sensitive parts of the object in the image with the target clothing material, the sensitive parts after filling with the target clothing material are integrated with the clothing material of the adjacent parts in the image into a whole. Since the integrity of the image is guaranteed, it is difficult to perceive that the image has been processed, thereby improving the user's viewing experience.

[0222] (2) By using a neural network model to determine the location of sensitive parts in an image, the accuracy of determining the location of sensitive parts can be improved, thereby improving the efficiency of subsequent filling of exposed sensitive parts.

[0223] (3) By using a neural network model to determine the tolerance based on the historical images viewed by the user, the accuracy of determining the tolerance can be improved. The image can be personalized based on the tolerance, which can not only save filling resources, but also improve the user's viewing experience.

[0224] (4) Determine whether to fill the image based on the exposed ratio of the object. While saving filling resources, the diversity of image filling can be achieved.

[0225] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for desensitizing images, characterized in that, The method includes: In response to an image display trigger operation, the image to be displayed is identified and processed; When an object in the image is identified to have exposed sensitive areas, the ratio of the area of ​​the exposed sensitive areas to the area of ​​the object is determined as the exposure ratio of the object. Determine the viewing user's tolerance level for the proportion of nudity of the subject; When the tolerance exceeds the tolerance threshold, a nudity ratio threshold corresponding to the tolerance threshold is determined, and target clothing material is filled into the sensitive area so that after the target clothing material is filled into the sensitive area, the nudity ratio of the object does not exceed the nudity ratio threshold. The target clothing material is used to blend with clothing materials in adjacent areas, where the adjacent areas are those connected to the sensitive areas. The image is displayed after the target clothing material has been filled.

2. The method according to claim 1, characterized in that, When the image to be displayed is a video frame from a video, before filling the sensitive area with target clothing material, the method further includes: When the similarity between the video frame and the reference frame exceeds a similarity threshold, clothing material filled in the sensitive parts of the reference frame is obtained, wherein the playback time of the reference frame is earlier than that of the video frame; The clothing material filled in the sensitive parts of the reference frame is used as the target clothing material for filling the sensitive parts of the video frame.

3. The method according to claim 1, characterized in that, When the image to be displayed is a video frame from a video, before filling the sensitive area with target clothing material, the method further includes: Determine the similarity between sensitive parts in the video frame and sensitive parts in the reference frame; When the similarity of the parts exceeds the similarity threshold, the clothing material filled in the sensitive parts of the reference frame will be used as the target clothing material to be filled in the sensitive parts of the video frame. The reference frame is played earlier than the video frame.

4. The method according to claim 1, characterized in that, Before filling the sensitive areas with target clothing material, the method further includes: Clothing material identification is performed on the adjacent parts; When it is identified that the adjacent area is filled with clothing material, the clothing material filled in the adjacent area is used as the target clothing material for filling the sensitive area of ​​the image.

5. The method according to claim 1, characterized in that, Before filling the sensitive areas with target clothing material, the method further includes: The clothing material of the object is identified; When it is identified that the object does not have clothing material, candidate clothing material belonging to the target clothing style is selected from the clothing material library as the target clothing material to be used to fill the object in the image; The target clothing material is used to fill the sensitive areas and the adjacent areas.

6. The method according to claim 5, characterized in that, The step of selecting candidate clothing materials belonging to the target clothing style from the clothing material library includes: Select candidate clothing materials from the clothing material library that meet the fusion conditions and are suitable for the body shape of the object; The fusion conditions include at least one of the following: The style of the candidate clothing material matches the style of the object's attire in the image; The type of the candidate clothing material is matched with the feature information of the object, wherein the feature information includes at least one of the object's occupation, gender, and personality. The types of candidate clothing materials are matched with the fill preferences of the viewing users.

7. The method according to claim 5, characterized in that, Before selecting candidate clothing materials belonging to the target clothing style from the clothing material library, the method further includes: The ratio of the exposed sensitive area to the area of ​​the object is determined as the exposure ratio of the object; When a user's tolerance for the nudity of the object exceeds the tolerance threshold, or when the nudity of the object exceeds the nudity ratio specified in the content review guidelines, it is determined that the operation of selecting candidate clothing materials belonging to the target clothing style from the clothing material library will be performed.

8. The method according to claim 1, characterized in that, The recognition processing of the image to be displayed includes: The first neural network model is invoked to perform the following processing: Extract feature vectors from multiple candidate regions in the image, and map each feature vector to the location of the candidate region including exposed sensitive parts and the confidence level. When the confidence level exceeds the confidence level threshold, it is determined that the candidate region includes exposed sensitive parts at the location. The first neural network model is trained based on a candidate sample region and labeled data for the candidate sample region. The labeled data includes the location and confidence level of the exposed sensitive parts in the candidate sample region.

9. The method according to claim 1, characterized in that, Determining the viewing user's tolerance for the proportion of nudity of the object includes: The second neural network model is invoked to perform the following processing: Extract the feature vectors of the sensitive parts and map the extracted feature vectors to tolerance. The second neural network model is trained using historical images viewed by the user and the user's tolerance for the proportion of exposed sensitive areas in the historical images as samples.

10. The method according to claim 1, characterized in that, Determining the viewing user's tolerance for the proportion of nudity of the object includes: Collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; The third neural network model is invoked to perform the following processing: Extract the feature vector of the behavior image, and map the extracted feature vector to tolerance; The third neural network model is trained using sample behavior images of the viewing users and the tolerance levels labeled for those sample behavior images.

11. The method according to claim 1, characterized in that, Determining the viewing user's tolerance for the proportion of nudity of the object includes: Collect behavioral images of the viewing user, wherein the behavioral images include at least one of facial expressions and body movements; The behavior image is matched with template images of different behavior types, and the behavior type corresponding to the successfully matched template image is taken as the behavior type of the viewing user. Based on the identified behavior type, the mapping relationship between different behavior types and tolerance is queried to obtain the tolerance corresponding to the identified behavior type.

12. An image desensitization processing apparatus, characterized in that, include: The recognition module is used to recognize and process the image to be displayed in response to the image display trigger operation; A filling module is used to determine the area ratio of the exposed sensitive parts to the object as the exposure ratio of the object when the object in the image is identified as having exposed sensitive parts; determine the viewing user's tolerance for the exposure ratio of the object; when the tolerance exceeds a tolerance threshold, determine the exposure ratio threshold corresponding to the tolerance threshold, and fill the sensitive parts with target clothing material so that after filling the sensitive parts with the target clothing material, the exposure ratio of the object does not exceed the exposure ratio threshold; The target clothing material is used to blend with clothing materials in adjacent areas, where the adjacent areas are those connected to the sensitive areas. The display module is used to display the image after the target clothing material has been filled.

13. The apparatus according to claim 12, characterized in that, When the image to be displayed is a video frame in a video, the filling module is further configured to, before filling the sensitive area with target clothing material, obtain the clothing material to be filled in the sensitive area of ​​the reference frame when the similarity between the video frame and the reference frame exceeds a similarity threshold, wherein the playback time of the reference frame is earlier than that of the video frame. The clothing material filled in the sensitive parts of the reference frame is used as the target clothing material for filling the sensitive parts of the video frame.

14. The apparatus according to claim 12, characterized in that, When the image to be displayed is a video frame in a video, the filling module is further configured to determine the part similarity between the sensitive part in the video frame and the sensitive part in the reference frame before filling the sensitive part with the target clothing material; When the similarity of the parts exceeds the similarity threshold, the clothing material filled in the sensitive parts of the reference frame will be used as the target clothing material to be filled in the sensitive parts of the video frame. The reference frame is played earlier than the video frame.

15. The apparatus according to claim 12, characterized in that, The filling module is also used to identify the clothing material in the adjacent parts before filling the sensitive parts with the target clothing material; When it is identified that the adjacent area is filled with clothing material, the clothing material filled in the adjacent area is used as the target clothing material for filling the sensitive area of ​​the image.

16. The apparatus according to claim 12, characterized in that, The filling module is also used to identify the clothing material of the object before filling the sensitive area with the target clothing material; When it is identified that the object does not have clothing material, candidate clothing material belonging to the target clothing style is selected from the clothing material library as the target clothing material to be used to fill the object in the image; The target clothing material is used to fill the sensitive areas and the adjacent areas.

17. The apparatus according to claim 16, characterized in that, The filling module is also used for: Select candidate clothing materials from the clothing material library that meet the fusion conditions and are suitable for the body shape of the object; The fusion conditions include at least one of the following: The style of the candidate clothing material matches the style of the object's attire in the image; The type of the candidate clothing material is matched with the feature information of the object, wherein the feature information includes at least one of the object's occupation, gender, and personality. The types of candidate clothing materials are matched with the fill preferences of the viewing users.

18. An electronic device, characterized in that, include: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the image desensitization processing method according to any one of claims 1 to 11.

19. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions, which, when executed, are used to implement the image desensitization processing method according to any one of claims 1 to 11.

20. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the image desensitization processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video processing method, device, electronic apparatus, and readable storage medium

    CN109040824A