Image data processing method, device, equipment and storage medium

The target image data is generated through face recognition and convolution processing, which solves the problem of privacy leakage when convolutional neural networks process image data in autonomous driving, improves the training effect of the model and maintains the integrity of face features.

CN114821726BActive Publication Date: 2025-05-06AUTOMOTIVE INTELLIGENCE & CONTROL OF CHINA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210463284.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-06
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

In autonomous driving technology, when convolutional neural networks are used to process large amounts of image data, it is easy to lead to the leakage of citizens' personal privacy. Existing solutions such as adding mosaics or directly blacking will cause image features to be missing, reducing the training effect of convolutional neural network models.

Method used

The target face image area in the image data is determined through face recognition, the target convolution kernel is generated and the convolution process is performed, the candidate convolution value greater than the target threshold is selected, the target area is determined, and the image mixing process is performed in the image data to generate the target image data.

Benefits of technology

On the premise of avoiding privacy leakage, the training effect of the convolutional neural network model should be improved as much as possible to ensure the feature integrity of the face image area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821726B_ABST
    Figure CN114821726B_ABST
Patent Text Reader

Abstract

The present application relates to an image data processing method, device, equipment and storage medium, and specifically to the field of computer vision technology. The method includes: obtaining and performing face recognition on image data, obtaining a target face image area in the image data; based on the pixel values ​​in the target face image area, generating a target convolution kernel and performing convolution processing on the image data to obtain each candidate convolution value; selecting at least one target convolution value greater than a target threshold, and determining a target area in the image data for generating the target convolution value; performing image mixing processing on the target area and the target face image area to obtain target image data. When training a neural network model based on the target image data generated by the above scheme, the training effect of the neural network model can be improved as much as possible while avoiding privacy leakage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an image data processing method, device, equipment and storage medium. Background Art

[0002] Autonomous driving or smart cars have a large number of sensors, and the amount of data generated is large and complex. Such a huge amount of data requires large-scale data processing and has higher requirements for data security.

[0003] In the existing autonomous driving technology, convolutional neural networks are needed as data processing algorithms to realize autonomous driving of smart cars. However, due to the use of convolutional neural networks, a large amount of image data is currently used in deep learning. During vehicle testing and data collection, a large amount of facial information will be collected on the roadside, which greatly threatens the personal privacy of citizens. In order to avoid excessive collection of facial information, the mainstream method is to annotate faces, add mosaics later, or directly blacken them to avoid the personal privacy of citizens in the training images of convolutional neural networks.

[0004] However, in the above scheme, adding mosaics or directly blackening the image will cause partial loss of image features, thereby reducing the training effect of the convolutional neural network model. Summary of the invention

[0005] The present application provides an image data processing method, apparatus, computer equipment and storage medium, which can improve the training effect of the neural network model as much as possible while avoiding privacy leakage. The technical solution is as follows.

[0006] In one aspect, a method for processing image data is provided, the method comprising:

[0007] Performing face recognition on the acquired image data to determine a target face image area in the image data;

[0008] Generate a target convolution kernel based on the pixel values ​​of the target face image area, and use the target convolution kernel to perform convolution processing on the image data to obtain various candidate convolution values;

[0009] Among the candidate convolution values, at least one target convolution value greater than a target threshold is selected, and a target area for generating the target convolution value is determined in the image data; the target area is an area other than the target face image area;

[0010] In the image data, the target area and the target face image area are subjected to image mixing processing to obtain target image data.

[0011] In another aspect, an image data processing device is provided, the device comprising:

[0012] A face recognition module, used to perform face recognition on the acquired image data and determine a target face image area in the image data;

[0013] An image convolution module, used to generate a target convolution kernel based on the pixel values ​​of the target face image area, and perform convolution processing on the image data using the target convolution kernel to obtain various candidate convolution values;

[0014] A target region determination module, configured to select at least one target convolution value greater than a target threshold value from among the candidate convolution values, and determine a target region for generating the target convolution value in the image data; the target region is a region other than the target face image region;

[0015] The image mixing module is used to perform image mixing processing on the target area and the target face image area in the image data to obtain target image data.

[0016] In a possible implementation, the image convolution module is further used to:

[0017] Get the face area size of the target face image area;

[0018] The target convolution kernel is generated by taking the size of the face region as the size of the convolution kernel and taking each pixel value of the target face image region as the weight in the convolution kernel.

[0019] In a possible implementation, the image convolution module is further used to:

[0020] When the face area size is larger than the target size threshold, the face area size is used as the size of the convolution kernel, and each pixel value of the target face image area is used as the weight in the convolution kernel to generate the target convolution kernel;

[0021] In a possible implementation manner, the device further includes:

[0022] The blur processing module is used to perform Gaussian blur processing on the target face image area in the image data when the size of the face area is less than or equal to the target size threshold, so as to obtain the target image data.

[0023] In a possible implementation, the device further includes:

[0024] The training image acquisition module is used to acquire the image data as the target image data when face recognition is performed on the image data and the recognition result indicates that there is no face image area in the image data.

[0025] In a possible implementation manner, the device further includes:

[0026] The target threshold determination module is used to sort the candidate convolution values ​​from large to small when the candidate convolution values ​​are obtained, and determine the candidate convolution value at the target position of the candidate convolution value sequence as the target threshold.

[0027] In a possible implementation, the image mixing module is further used to:

[0028] In the image data, each pixel value in the target area is nonlinearly mixed with each pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

[0029] In a possible implementation, the image mixing module is further used to:

[0030] Randomly select half of the pixel values ​​in the target area and half of the pixel values ​​in the target face image area for splicing to obtain a spliced ​​pixel area;

[0031] In the image data, each pixel value in the spliced ​​pixel area is nonlinearly mixed with the pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

[0032] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned image data processing method.

[0033] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned image data processing method.

[0034] In another aspect, a computer program product or a computer program is provided, wherein the computer program product or the computer program comprises computer instructions, wherein the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned image data processing method.

[0035] The technical solution provided by this application may have the following beneficial effects:

[0036] After acquiring the image data, the computer device may first perform face recognition on the image data, and after identifying the face image area, a target convolution kernel may be generated according to the pixel value of the face image area to perform convolution processing on the image data; at this time, when the obtained candidate convolution value is large, it means that the area corresponding to the candidate convolution value has a high feature similarity with the face image area, and therefore a larger target convolution value may be selected from the candidate convolution values, and a target area for generating the target convolution value may be selected from the image data, and at this time the target area is outside the face image area; therefore, after the target area is mixed with the face image area in the acquired image data, the generated target image data can ensure the integrity of the features of the face image area as much as possible while avoiding privacy leakage caused by the face image, and when the neural network model is trained through the target image data, the training effect of the neural network model can be improved as much as possible while avoiding privacy leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 It is a structural schematic diagram of an image data processing system according to an exemplary embodiment.

[0039] Figure 2 The figure is a flowchart of a method for processing image data according to an exemplary embodiment.

[0040] Figure 3 The figure is a flowchart of a method for processing image data according to an exemplary embodiment.

[0041] Figure 4 A schematic diagram of image data acquisition involved in an embodiment of the present application is shown.

[0042] Figure 5 It is a structural block diagram of an image data processing device according to an exemplary embodiment.

[0043] Figure 6 A structural block diagram of a computer device shown in an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0044] The technical solution of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0045] It should be understood that the "indication" mentioned in the embodiments of the present application can be a direct indication, an indirect indication, or an indication of an association relationship. For example, A indicates B, which can mean that A directly indicates B, for example, B can be obtained through A; it can also mean that A indirectly indicates B, for example, A indicates C, and B can be obtained through C; it can also mean that there is an association relationship between A and B.

[0046] In the description of the embodiments of the present application, the term "corresponding" may indicate a direct or indirect correspondence between two items, or an association relationship between the two items, or a relationship between indication and being indicated, configuration and being configured, and the like.

[0047] In an embodiment of the present application, "predefinition" can be achieved by pre-saving corresponding codes, tables or other methods that can be used to indicate relevant information in a device (for example, including a terminal device and a network device). The present application does not limit its specific implementation method.

[0048] Figure 1 1 is a schematic diagram of a structure of an image data processing system according to an exemplary embodiment. The image data processing system includes a server 110 and a target vehicle 120. The target vehicle 120 may include modules such as a data processing device, an image acquisition device, and a data storage module.

[0049] Optionally, the target vehicle 120 includes an image acquisition device and a data storage module. The image acquisition device can acquire images of the environment around the target vehicle during the operation of the target vehicle, and store the acquired images in the data storage module in the target vehicle.

[0050] Optionally, the target vehicle 120 is communicated with the server 110 via a transmission network (such as a wireless communication network). The target vehicle 120 can upload various data (such as collected images) stored in the data storage module to the server 110 via the wireless communication network, so that the server 110 processes the collected images and trains convolutional neural network models used in intelligent driving and other aspects based on the collected images.

[0051] Optionally, the target vehicle 120 also includes a state acquisition device ( Figure 1(not shown), the state acquisition device can collect the driving state of the target vehicle 120 during driving in real time, and save the driving state as time series data in the data storage device in the target vehicle 120. At this time, the target vehicle can also upload the driving state in the data storage module to the server 110 through the wireless communication network, so that the server 110 can train the convolutional neural network model applied to intelligent driving and other aspects according to the image collected by the target vehicle and the driving state of the target vehicle.

[0052] Optionally, the target vehicle also includes a data processing device. When the image acquisition device of the target vehicle 120 captures the image, the data processing device can recognize the image and determine whether there is facial information in the captured image. When facial information is detected, the data processing device can perform privacy processing on the captured image, and mask the facial information in the captured image through an image processing method before uploading it to the server 110, so as to prevent the captured image from infringing on the privacy of others.

[0053] Optionally, when receiving the image captured by the target vehicle 120, the server 110 may also recognize the image to determine whether there is facial information in the captured image. When facial information is detected, the server may perform privacy processing on the captured image, mask the facial information in the captured image through an image processing method, and then save the image into the training data set of the convolutional neural network model to avoid the convolutional neural network model's training data set containing other people's privacy information.

[0054] Optionally, the server 110 can also establish wireless communication connections with various new energy vehicles (for example, including the target vehicle 120) that establish communication connections with the server 110, including the target vehicle 120, through a wireless communication network, and send corresponding algorithm information to each new energy vehicle. For example, the server 110 can send the model parameters of the trained convolutional neural network model to the target vehicle 120 through the wireless communication network. At this time, the intelligent driving application in the target vehicle 120 can load the trained convolutional neural network model to realize real-time processing of the real-time collected images, and determine the driving mode of the target vehicle at this time, thereby realizing intelligent driving.

[0055] Optionally, the above-mentioned server can be a server cluster or a distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms and other technical computing services.

[0056] Optionally, the system may further include a management device, which is used to manage the system (such as managing the connection status between each module and the server, etc.), and the management device is connected to the server via a communication network. Optionally, the communication network is a wired network or a wireless network.

[0057] Optionally, the above-mentioned wireless network or wired network uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any other network, including but not limited to any combination of local area network, metropolitan area network, wide area network, mobile, limited or wireless network, private network or virtual private network. In some embodiments, the technology and / or format including hypertext markup language, extensible markup language, etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as secure socket layer, transport layer security, virtual private network, Internet protocol security, etc. can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technology can also be used to replace or supplement the above-mentioned data communication technology.

[0058] A convolutional neural network model may be used in the intelligent driving process involved in the target vehicle. In order to train the convolutional neural network model, an image acquisition device can be deployed on the target vehicle. During the driving process of the target vehicle, the image acquisition device on the target vehicle collects the environmental information around the target vehicle in real time (for example, collecting road information during the driving process of the target vehicle, etc.), and then annotates the collected images and uses them as training images for the convolutional neural network.

[0059] However, during the driving process of the target vehicle, there may be pedestrians on the road. At this time, the image acquisition device on the target vehicle is likely to collect the facial information of pedestrians on the road. At this time, if the image containing facial information is directly transmitted to the server as a training image for training the convolutional neural network model, it is easy to cause infringement of citizens' privacy. Therefore, the image data obtained by the image acquisition device cannot be directly uploaded to the server as a training image. The computer equipment needs to perform certain processing on the collected images to avoid infringing on citizens' privacy while ensuring the training effect of the collected images as training images.

[0060] Figure 2 is a flowchart of an image data processing method according to an exemplary embodiment. The method is executed by a computer device, which may be Figure 1 The data processing device in the target vehicle shown in FIG. Figure 2 As shown, the image data processing method may include the following steps:

[0061] Step 201 , performing face recognition on the acquired image data to determine a target face image region in the image data.

[0062] After acquiring the image data collected by the image acquisition device of the target vehicle, the computer device can first perform face recognition processing on the image data, for example, through a target detection algorithm, etc., to perform face recognition on the image data to determine whether there is a face image area in the image data.

[0063] When there is no face image area in the image data, the computer device skips the image data and processes the next image data; when there is a face image area in the image data, the computer device can generate a corresponding detection frame through a target detection algorithm or other means to determine the position of the target face image area in the image data.

[0064] Step 202: Generate a target convolution kernel based on the pixel values ​​of the target face image area, and use the target convolution kernel to perform convolution processing on the image data to obtain various candidate convolution values.

[0065] When a target facial image area of ​​the image data is detected, the computer device can generate a target convolution kernel based on the pixel values ​​in the target facial image area. For example, the computer device can directly use the pixel values ​​of the RGB (Red Green Blue) three channels in the target facial image area as the weights of the convolution kernels of the three channels, respectively, to generate a three-channel target convolution kernel of the same size as the target facial image area; or the computer device can use the weighted sum of the RGB three channels of the pixel values ​​in the target facial image area as the weight of the convolution kernel, to generate a single-channel target convolution kernel of the same size as the target facial image area; or in an embodiment of the present application, other methods can be used to process the pixel values ​​in the target facial image area to obtain weights, thereby generating a target convolution kernel of the same size as the target facial image area, so as to perform convolution processing on the image data.

[0066] The target convolution kernel generated by the above scheme has weights in the convolution kernel that are generated based on the pixel values ​​in the target face image area. Therefore, in the image data, the closer the pixel features are to the sub-region of the target face image area, the larger the convolution value obtained after convolution processing by the convolution kernel. Therefore, the size of each candidate convolution value represents the degree of similarity between each sub-region and the target face image area in the image data.

[0067] Step 203, selecting at least one target convolution value greater than a target threshold value from among the candidate convolution values, and determining a target region for generating the target convolution value in the image data.

[0068] The target area is an area other than the target face image area.

[0069] Since in the embodiment of the present application, the size of each candidate convolution value represents the similarity between each sub-region in the image data and the target face image region. Therefore, when each candidate convolution value is obtained, the computer device can select at least one target convolution value greater than the target threshold value from each candidate convolution value, and determine the target region for generating the target convolution value in the image data, that is, select at least one region in the image data outside the target face image region, which has pixel features similar to the target face image region.

[0070] Step 204, in the image data, the target area and the target face image area are subjected to image mixing processing to obtain target image data for training the neural network model.

[0071] After the target area is selected, the computer device can perform image blending processing on the pixels of the target face image area and the pixels of the target area. For example, the pixels of the non-face area (that is, the target area) are mixed with the pixels of the target face image area by image linear blending or image nonlinear blending. At this time, the pixels of the target face image area are disturbed and the facial features are blocked, thereby reducing the risk of privacy infringement to a certain extent. In addition, since the pixel features of the non-face area are similar to the pixel features of the target face image area, when the image data obtained after the blending process is used as the target image data, the training effect of the neural network model is also guaranteed to a certain extent.

[0072] Optionally, when there are multiple faces in the image data, after performing face recognition on the image data, multiple target face image areas can be detected in the image data.

[0073] At this time, each target face image region may be processed by the method shown in step 202 to step 204, so that each target face image region is subjected to image mixing, and finally the image data is updated to the target image data.

[0074] In summary, after acquiring the image data, the computer device can first perform face recognition on the image data, and after identifying the face image area, a target convolution kernel can be generated according to the pixel value of the face image area to perform convolution processing on the image data; at this time, when the obtained candidate convolution value is large, it means that the area corresponding to the candidate convolution value has a high feature similarity with the face image area, so a larger target convolution value can be selected from the candidate convolution values, and a target area for generating the target convolution value can be selected from the image data, and at this time the target area is outside the face image area; therefore, after the target area is mixed with the face image area in the acquired image data, the generated target image data can ensure the integrity of the features of the face image area as much as possible under the premise of avoiding privacy leakage caused by the face image, and when the neural network model is trained through the target image data, the training effect of the neural network model can be improved as much as possible under the premise of avoiding privacy leakage.

[0075] Figure 3 is a method flow chart of an image data processing method according to an exemplary embodiment. The method is executed by a computer device, which may be Figure 1 The data processing device in the target vehicle shown in FIG. Figure 3 As shown, the image data processing method may include the following steps:

[0076] Step 301: perform face recognition on the image data to determine a target face image area in the image data.

[0077] In an embodiment of the present application, after the computer device acquires the image data, it can perform face recognition on the image data through a target detection algorithm.

[0078] Optionally, the target detection algorithm can be a target detection algorithm such as yolo SSD that can realize face recognition to generate a face detection frame. The format of the detection frame is (center point coordinates, detection frame height, detection frame width). At this time, the computer device can determine the size of the target face image area by obtaining the width and height of the detection frame.

[0079] In a possible implementation, when face recognition is performed on the image data and the recognition result indicates that there is no face image area in the image data, the image data is acquired as the target image data.

[0080] That is, when a computer device performs face recognition on image data through a target detection algorithm and does not detect a face image area in the image data, it may be that there is no face image in the image data, or the face image in the image data is very small and cannot be recognized by the target detection algorithm; but whether there is no face image or the face image is very small and cannot be recognized, it can be considered that it will not cause the leakage of face information. At this time, the image data can be directly obtained as the target image data to train the convolutional neural network.

[0081] Step 302, obtaining the face area size of the target face image area.

[0082] When the computer device performs face recognition on the image data through the target detection algorithm and detects the target face image area in the target image data, the computer device can first determine the face area size of the target face image area, such as determining the pixel area of ​​the target face image area, that is, the product of the pixel width and pixel height of the target face image area.

[0083] In a possible implementation, when the size of the face region is less than or equal to a target size threshold, Gaussian blur processing is performed on the target face image region in the image data to obtain the target image data.

[0084] When the product of the pixel width and pixel height of the target face image area is smaller than the target size threshold, for example, smaller than 25 pixels, the pixel information of the target face image area is relatively small. At this time, the computer device can perform a simple Gaussian blur processing on the target face image area, that is, it can adequately cover the face without excessively interfering with the pixel features of the target face image area.

[0085] Among them, the two-dimensional Gaussian distribution can be expressed by the following formula:

[0086]

[0087] In an embodiment of the present application, by adjusting a relatively large blur value, the pixels of the target facial image area can be processed using the above formula (i.e., the pixel values ​​of each pixel point in the target facial image area are transformed using the above two-dimensional Gaussian distribution formula), and the important facial features can be blurred and transmitted to the cloud server as target image data.

[0088] Step 303, using the face area size as the size of the convolution kernel, and using each pixel value in the target face image area as the weight in the convolution kernel, to generate the target convolution kernel.

[0089] In one possible implementation, when the face area size is larger than the target size threshold, the face area size is used as the size of the convolution kernel, and each pixel value of the target face image area is used as the weight in the convolution kernel to generate the target convolution kernel.

[0090] In an embodiment of the present application, when the face area size is larger than the target size threshold, simple blurring (such as Gaussian blurring) is not sufficient to cover up important face information. At this time, the computer device can use the face area size as the size of the convolution kernel, and use the pixel values ​​of the target face image area as the weights in the convolution kernel to generate a target convolution kernel.

[0091] For example, when the target face image area contains pixel values ​​of the three RGB channels, the convolution kernels corresponding to the width and height can be generated according to the pixel values ​​of the three RGB channels respectively, and corresponding to the three channels (that is, the convolution kernel with a width of w times a height of h times 3).

[0092] Step 304: perform convolution processing on the image data using the target convolution kernel to obtain various candidate convolution values.

[0093] In the embodiment of the present application, after the target convolution kernel is obtained, a specified step size (such as a step size of 1) can be set to perform convolution processing on the image data, thereby obtaining candidate convolution values ​​corresponding to each sub-region in the target convolution kernel. Since the target convolution kernel is obtained based on the pixel value of the target face image region, the convolution value obtained by convolving the target face image region with the target convolution kernel is theoretically the largest. In addition, the larger the convolution value obtained by convolving the target convolution kernel, the greater the pixel similarity between the sub-region in the target image region and the target face image region.

[0094] In a possible implementation, a target convolution kernel is used to perform convolution processing on regions in the image data except for the target face image region, thereby obtaining various candidate convolution values.

[0095] Since the convolution is performed through the target convolution kernel, what needs to be screened out is the area close to the target face image area. Therefore, when the image data is convolved through the target convolution kernel, the target face image area can be directly excluded, thereby reducing the computing resources consumed by the computer during convolution processing and improving computing efficiency.

[0096] Step 305 , selecting at least one target convolution value greater than a target threshold value from among the candidate convolution values, and determining a target region for generating the target convolution value in the image data.

[0097] The target area is an area other than the target face image area.

[0098] In a possible implementation, when the candidate convolution values ​​are obtained, the candidate convolution values ​​are sorted from large to small, and the candidate convolution value located at a target position in the candidate convolution value sequence is determined as the target threshold.

[0099] For example, after the computer device obtains each candidate convolution value, it can sort the candidate convolution values ​​from large to small, and determine the candidate convolution value at the eleventh position as the target threshold. At this time, the computer device can select at least one target convolution value from the maximum ten candidate convolution values ​​and determine the target area corresponding to the target convolution value.

[0100] Alternatively, in a possible implementation, after obtaining each candidate convolution value, the computer device may determine the largest one among the candidate convolution values ​​as the target convolution value, and obtain the target area corresponding to the target convolution value.

[0101] Step 306: In the image data, perform image mixing processing on the target area and the target face image area to obtain target image data.

[0102] Taking the selection of a target area corresponding to a target convolution value as an example, in a possible implementation method, in the image data, each pixel value in the target area is nonlinearly mixed with each pixel value in the target face image area, and the mixed pixel values ​​are updated to the pixel values ​​of the target face image area to obtain the target image data.

[0103] Alternatively, after a plurality of target convolution values ​​are selected, the target regions corresponding to the respective target convolution values ​​may be image mixed with the target region one by one, thereby obtaining final target image data.

[0104] In a possible implementation, half of the pixel values ​​in the target area and half of the pixel values ​​in the target face image area are randomly selected for splicing to obtain a spliced ​​pixel area;

[0105] In the image data, each pixel value in the spliced ​​pixel area is nonlinearly mixed with the pixel value in the target face image area, and the mixed pixel values ​​are updated to the pixel values ​​of the target face image area to obtain the target image data.

[0106] For example, a computer device can randomly select half of the pixel values ​​in the target area and randomly select the other half of the pixel values ​​in the target face image area for splicing to obtain a spliced ​​pixel area, and then nonlinearly mix the spliced ​​pixel area with the pixel values ​​in the target face image area. At this time, the proportion of pixels in the target face image area can be further reduced.

[0107] Alternatively, in another possible implementation, after multiple target convolution values ​​are selected, part of the pixel values ​​can be randomly selected from the target areas corresponding to the multiple target convolution values ​​and then spliced ​​together, and then the spliced ​​mixed pixel area is nonlinearly mixed with the target face image area, so that the updated target face image area has more overall image features.

[0108] Please refer to Figure 4 , which shows a schematic diagram of image data acquisition involved in the embodiment of the present application. Figure 4 As shown in the figure, after the overall autonomous driving data collection module of the target vehicle is turned on, the face recognition function is turned on synchronously. The currently collected image data (that is, pictures) are judged, and pictures without faces are directly uploaded to the local hard disk or cloud. Pictures with faces are recognized, and they need to be processed later, so as not to affect the processing of subsequent data. Pictures with faces entering the big data buffer need to go through the next step of judgment.

[0109] First, if the face box is smaller than 25 pixels, the data is labeled as critical but not sensitive. Since the information is relatively small, it only needs to be simply blurred, and Gaussian blur can be used to adequately block the face.

[0110] Adjust a relatively large blur value to blur the important features of the face, and then transfer it to the local hard drive or cloud.

[0111] Finally, when the facial contour and details are relatively clear, they are defined as critical and sensitive information and input into subsequent processing. As mentioned above, if the area around the face is blackened or similar operations are performed, noise will appear in the image. For different models, this will cause greater or lesser impacts. Therefore, blurring the face alone may not meet the information security benchmark, and drastically modifying the image will affect performance. Therefore, this application uses a method of fusing the non-face area and the face frame in the current image to ensure the quality of the data with minimal impact on the image quality. The modification steps are shown in step 306 and will not be repeated here.

[0112] The above method can be used to shield the face information while protecting the overall smoothness of the original image, and will not affect the overall collection efficiency.

[0113] Finally, the processed data is transferred to the local hard disk or cloud. At this point, all the data is saved after different processing.

[0114] In summary, after acquiring the image data, the computer device can first perform face recognition on the image data, and after identifying the face image area, a target convolution kernel can be generated according to the pixel value of the face image area to perform convolution processing on the image data; at this time, when the obtained candidate convolution value is large, it means that the area corresponding to the candidate convolution value has a high feature similarity with the face image area, so a larger target convolution value can be selected from the candidate convolution values, and a target area for generating the target convolution value can be selected from the image data, and at this time the target area is outside the face image area; therefore, after the target area is mixed with the face image area in the acquired image data, the generated target image data can ensure the integrity of the features of the face image area as much as possible under the premise of avoiding privacy leakage caused by the face image, and when the neural network model is trained through the target image data, the training effect of the neural network model can be improved as much as possible under the premise of avoiding privacy leakage.

[0115] Figure 5 1 is a block diagram showing a structure of an image data processing device according to an exemplary embodiment. The image data processing device includes:

[0116] A face recognition module 501 is used to perform face recognition on the acquired image data and determine a target face image area in the image data;

[0117] An image convolution module 502 is used to generate a target convolution kernel based on the pixel values ​​of the target face image area, and perform convolution processing on the image data using the target convolution kernel to obtain various candidate convolution values;

[0118] A target region determination module 503 is used to select at least one target convolution value greater than a target threshold value from the candidate convolution values, and determine a target region for generating the target convolution value in the image data; the target region is a region other than the target face image region;

[0119] The image mixing module 504 is used to perform image mixing processing on the target area and the target face image area in the image data to obtain target image data.

[0120] In a possible implementation, the image convolution module is further used to:

[0121] Get the face area size of the target face image area;

[0122] The target convolution kernel is generated by taking the size of the face region as the size of the convolution kernel and taking each pixel value of the target face image region as the weight in the convolution kernel.

[0123] In a possible implementation, the image convolution module is further used to:

[0124] When the face area size is larger than the target size threshold, the face area size is used as the size of the convolution kernel, and each pixel value of the target face image area is used as the weight in the convolution kernel to generate the target convolution kernel;

[0125] In a possible implementation manner, the device further includes:

[0126] The blur processing module is used to perform Gaussian blur processing on the target face image area in the image data when the size of the face area is less than or equal to the target size threshold, so as to obtain the target image data.

[0127] In a possible implementation, the device further includes:

[0128] The training image acquisition module is used to acquire the image data as the target image data when face recognition is performed on the image data and the recognition result indicates that there is no face image area in the image data.

[0129] In a possible implementation manner, the device further includes:

[0130] The target threshold determination module is used to sort the candidate convolution values ​​from large to small when the candidate convolution values ​​are obtained, and determine the candidate convolution value at the target position of the candidate convolution value sequence as the target threshold.

[0131] In a possible implementation, the image mixing module is further used to:

[0132] In the image data, each pixel value in the target area is nonlinearly mixed with each pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

[0133] In a possible implementation, the image mixing module is further used to:

[0134] Randomly select half of the pixel values ​​in the target area and half of the pixel values ​​in the target face image area for splicing to obtain a spliced ​​pixel area;

[0135] In the image data, each pixel value in the spliced ​​pixel area is nonlinearly mixed with the pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

[0136] In summary, after acquiring the image data, the computer device can first perform face recognition on the image data, and after identifying the face image area, a target convolution kernel can be generated according to the pixel value of the face image area to perform convolution processing on the image data; at this time, when the obtained candidate convolution value is large, it means that the area corresponding to the candidate convolution value has a high feature similarity with the face image area, so a larger target convolution value can be selected from the candidate convolution values, and a target area for generating the target convolution value can be selected from the image data, and at this time the target area is outside the face image area; therefore, after the target area is mixed with the face image area in the acquired image data, the generated target image data can ensure the integrity of the features of the face image area as much as possible under the premise of avoiding privacy leakage caused by the face image, and when the neural network model is trained through the target image data, the training effect of the neural network model can be improved as much as possible under the premise of avoiding privacy leakage.

[0137] Figure 6 The structural block diagram of a computer device 600 shown in an exemplary embodiment of the present application is shown. The computer device can be implemented as a server in the above-mentioned solution of the present application. The computer device 600 includes a central processing unit (CPU) 601, a system memory 604 including a random access memory (RAM) 602 and a read-only memory (ROM) 603, and a system bus 605 connecting the system memory 604 and the central processing unit 601. The computer device 600 also includes a mass storage device 606 for storing an operating system 609, an application program 610 and other program modules 611.

[0138] The mass storage device 606 is connected to the central processing unit 601 via a mass storage controller (not shown) connected to the system bus 605. The mass storage device 606 and its associated computer readable medium provide non-volatile storage for the computer device 600. That is, the mass storage device 606 may include a computer readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0139] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, Erasable Programmable Read Only Memory (EPROM), Electronically-Erasable Programmable Read-Only Memory (EEPROM) flash memory or other solid-state storage technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, cassettes, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above. The above-mentioned system memory 604 and mass storage device 606 can be collectively referred to as memory.

[0140] According to various embodiments of the present disclosure, the computer device 600 can also be connected to a remote computer on the network through a network such as the Internet. That is, the computer device 600 can be connected to the network 608 through the network interface unit 607 connected to the system bus 605, or the network interface unit 607 can be used to connect to other types of networks or remote computer systems (not shown).

[0141] The memory further includes at least one computer program, which is stored in the memory. The central processing unit 601 implements all or part of the steps in the methods shown in the above embodiments by executing the at least one computer program.

[0142] In an exemplary embodiment, a computer-readable storage medium is also provided, which is used to store at least one computer program, and the at least one computer program is loaded and executed by a processor to implement all or part of the steps in the above method. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0143] In an exemplary embodiment, a computer program product or a computer program is also provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned Figure 2 or Figure 3 All or part of the steps of the method shown in any embodiment.

[0144] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0145] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for processing image data, characterized in that: The method comprises: Performing face recognition on the acquired image data to determine a target face image area in the image data; Generate a target convolution kernel based on the pixel values ​​of the target face image area, and use the target convolution kernel to perform convolution processing on the image data to obtain various candidate convolution values; Selecting at least one target convolution value greater than a target threshold value from among the candidate convolution values, and determining a target region for generating the target convolution value in a region of the image data other than the target face image region; In the image data, the target area and the target face image area are subjected to image mixing processing to obtain target image data.

2. The method according to claim 1, characterized in that The generating a target convolution kernel based on the pixel values ​​of the target face image area includes: Get the face area size of the target face image area; The target convolution kernel is generated by taking the size of the face region as the size of the convolution kernel and taking each pixel value of the target face image region as the weight in the convolution kernel.

3. The method according to claim 2, characterized in that The method of using the face area size as the size of the convolution kernel and each pixel value of the target face image area as the weight in the convolution kernel to generate the target convolution kernel includes: When the face area size is larger than the target size threshold, the face area size is used as the size of the convolution kernel, and each pixel value of the target face image area is used as the weight in the convolution kernel to generate the target convolution kernel; The method further comprises: When the size of the face area is less than or equal to the target size threshold, Gaussian blur processing is performed on the target face image area in the image data to obtain the target image data.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: When face recognition is performed on the image data and the recognition result indicates that there is no face image area in the image data, the image data is acquired as the target image data.

5. The method according to any one of claims 1 to 3, characterized in that: Before selecting at least one target convolution value greater than a target threshold value from among the candidate convolution values, the method further includes: When the candidate convolution values ​​are obtained, the candidate convolution values ​​are sorted from large to small, and the candidate convolution value located at the target position of the candidate convolution value sequence is determined as the target threshold.

6. The method according to any one of claims 1 to 3, characterized in that: The step of performing image mixing processing on the target area and the target face image area in the image data to obtain the target image data comprises: In the image data, each pixel value in the target area is nonlinearly mixed with each pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

7. The method according to any one of claims 1 to 3, characterized in that: The step of performing image mixing processing on the target area and the target face image area in the image data to obtain the target image data comprises: Randomly select half of the pixel values ​​in the target area and half of the pixel values ​​in the target face image area for splicing to obtain a spliced ​​pixel area; In the image data, each pixel value in the spliced ​​pixel area is nonlinearly mixed with the pixel value in the target face image area, and each pixel value obtained by mixing is updated to the pixel value of the target face image area to obtain the target image data.

8. An image data processing device, characterized in that: The device comprises: A face recognition module, used to perform face recognition on the acquired image data and determine a target face image area in the image data; An image convolution module, used to generate a target convolution kernel based on the pixel values ​​of the target face image area, and perform convolution processing on the image data using the target convolution kernel to obtain various candidate convolution values; A target region determination module, configured to select at least one target convolution value greater than a target threshold value from among the candidate convolution values, and determine a target region for generating the target convolution value in a region of the image data other than the target face image region; the target region is a region other than the target face image region; The image mixing module is used to perform image mixing processing on the target area and the target face image area in the image data to obtain target image data.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the image data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the image data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Face encryption method and device, face recognition method and device, electronic equipment and storage medium

    CN111931145A

  • Face recognition method and device, equipment and medium

    CN113762033A