Image data processing method, device, electronic device and readable storage medium

By extracting background images from user images and performing feature extraction and compression processing, the problem that image feature extraction in the prior art is difficult to combine local and global information, improving the accuracy of image retrieval and reducing computing and memory requirements.

CN113515659BActive Publication Date: 2025-05-06PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110470828.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-28
Publication Date
2025-05-06
Estimated Expiration
2041-04-28

AI Technical Summary

Technical Problem

When extracting image features in the prior art, it is difficult to effectively combine local information and global information in the image, resulting in low accuracy of image retrieval results, and large calculation amounts and large memory usage.

Method used

By extracting the background image from the user image, then extracting the deep feature map from the background image, and dividing it into a preset number of feature map sub-blocks, inputting points to feature networks to obtain the first image feature, and compressing it to obtain the second image feature.

Benefits of technology

It improves the accuracy of image retrieval, reduces the computational amount of data processing and memory space usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113515659B_ABST
    Figure CN113515659B_ABST
Patent Text Reader

Abstract

The present invention relates to image processing, and discloses an image data processing method, including: obtaining a user image to be processed, wherein the user image includes a foreground portrait and a background image; extracting a background image from the user image; extracting a first image feature from the background image, specifically including: extracting a deep feature map from the background image, dividing the deep feature map into a preset number of feature map sub-blocks, and inputting the preset number of feature map sub-blocks into a point pair feature network to obtain the first image feature; compressing the first image feature to obtain a second image feature. The present invention also provides an image data processing device, an electronic device, and a readable storage medium. The present invention extracts image features that combine local information and global information from an image, which can improve the accuracy of subsequent image retrieval using image features, and reduce the amount of calculation and memory space occupied by data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image data processing method, device, electronic device and readable storage medium. Background Art

[0002] Content-based image retrieval technology has been well applied in many fields. One typical application is the field of financial fraud detection. For example, in video interview scenarios, backgrounds with higher fraud risks (commonly known as black backgrounds, usually in fixed places such as hotels) can be retrieved.

[0003] Content-based image retrieval usually performs retrieval based on image features, so an important step in processing image data is to extract image features. Taking video face-to-face review as an example, the user image requesting the face-to-face review is generally taken by a mobile terminal. Due to the incomplete fixation of the mobile device, the user's portrait and background information will have a large degree of scale transformation, occlusion, and missing problems. At present, some conventional deep learning algorithms learn the global information or foreground target information of the entire image in the process of extracting image features. Since the data is processed by point cloud information (pixel level), the data processing involves a large amount of calculation and memory usage. In addition, the aforementioned method fails to combine the local information and global information in the image well, so the extracted image features may affect the accuracy of the image retrieval results. Summary of the invention

[0004] In view of the above, it is necessary to provide an image data processing method, which aims to extract image features that combine local information and global information from images, improve the accuracy of subsequent image retrieval using image features, and reduce the amount of computation and memory space occupied by data processing.

[0005] An embodiment of the present invention provides an image data processing method, the method comprising:

[0006] Acquire a user image to be processed, wherein the user image includes a foreground portrait and a background image;

[0007] Extracting a background image from the user image;

[0008] Extracting a first image feature from the background image specifically includes: extracting a deep feature map from the background image, dividing the deep feature map into a preset number of feature map sub-blocks, and inputting the preset number of feature map sub-blocks into a point pair feature network to obtain the first image feature;

[0009] The first image feature is compressed to obtain a second image feature.

[0010] Optionally, the step of extracting a background image from the user image includes:

[0011] Detecting a foreground person image and a background image from the user image using a pre-trained segmentation model;

[0012] The pixel values ​​of the foreground portrait area are set to 0, the pixel values ​​of the background image area are set to 1, and a first two-dimensional matrix having the same size as the user image is output;

[0013] The element values ​​of the first two-dimensional matrix are multiplied by the pixel values ​​of corresponding pixels in the user image to obtain a background image with the foreground portrait subtracted.

[0014] Optionally, the step of extracting a deep feature map from the background image includes:

[0015] Image features are extracted from the background image to obtain a H*W*C deep feature map, where H, W, and C are the height, width, and number of the deep feature map, respectively.

[0016] Optionally, dividing the deep feature map into a preset number of feature map sub-blocks includes:

[0017] The H*W*C deep feature map is evenly divided into M*N feature map sub-blocks with C layers through an M*N grid of a preset size.

[0018] Optionally, the aspect ratio M / N of the M*N grid is equal to the aspect ratio of the background image.

[0019] Optionally, inputting the preset number of feature map sub-blocks into a point pair feature network to obtain the first image feature includes:

[0020] M*N feature map sub-blocks with C layers are used as input of the point-to-point feature network, and the point-to-point feature network outputs M*N one-dimensional local features based on global information with a preset D dimension, that is, M*N*D-dimensional features, as the first image features.

[0021] Optionally, compressing the first image feature to obtain the second image feature includes:

[0022] Scaling the first two-dimensional matrix in equal proportion to the size of the feature map sub-block to obtain a second two-dimensional matrix;

[0023] Multiplying the second two-dimensional matrix by the first image feature to obtain a new M*N*D dimensional feature;

[0024] The new M*N*D-dimensional feature is input into a pooling layer external to the point pair feature network and converted into a D-dimensional one-dimensional vector as the second image feature of the background image.

[0025] An embodiment of the present invention further provides an image data processing device, the device comprising:

[0026] An image acquisition module, used to acquire a user image to be processed, wherein the user image includes a foreground portrait and a background image;

[0027] A background separation module, used to extract a background image from the user image;

[0028] A feature extraction module is used to extract the first image feature from the background image, specifically comprising: extracting a deep feature map from the background image; dividing the deep feature map into a preset number of feature map sub-blocks; inputting the preset number of feature map sub-blocks into a point pair feature network to obtain the first image feature;

[0029] The feature processing module is used to compress the first image feature to obtain a second image feature.

[0030] An embodiment of the present invention further provides an electronic device, the electronic device comprising:

[0031] at least one processor; and,

[0032] a memory communicatively connected to the at least one processor; wherein,

[0033] The memory stores an image data processing program executable by the at least one processor, and the image data processing program is executed by the at least one processor so that the at least one processor can perform the image data processing method described above.

[0034] An embodiment of the present invention also provides a computer-readable storage medium, on which an image data processing program is stored. The image data processing program can be executed by one or more processors to implement the image data processing method as described above.

[0035] Compared with the prior art, the image data processing method provided by the present invention obtains a user image to be processed, extracts a background image from the user image, and then extracts a first image feature from the user image, and finally compresses the first image feature to obtain a second image feature that can be used to perform image retrieval. In the step of extracting the first image feature, specifically, a deep feature map is first extracted from the background image, and then the deep feature map is divided into a preset number of feature map sub-blocks, and then the preset number of feature map sub-blocks are input into a point pair feature network, and the first image feature is output. This solves the problem in the existing feature extraction method that the extracted features cannot well combine the local information and the global information in the image due to the large degree of scale transformation, occlusion, and missing of the user portrait and background information, which ultimately affects the accuracy of the image retrieval results. It also reduces the amount of calculation and saves memory space. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A schematic diagram of an application environment provided by an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of the hardware architecture of an electronic device provided by an embodiment of the present invention;

[0038] Figure 3 A schematic diagram of functional modules of an image data processing device provided by an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of a flow chart of an image data processing method provided by an embodiment of the present invention;

[0040] Figure 5 FIG. 1 is a schematic diagram of a background image with a foreground portrait subtracted according to an embodiment of the present invention.

[0041] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0043] It should be noted that the descriptions of "first", "second", etc. in the present invention are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0044] See also Figure 1 The figure shows an optional application environment diagram of the first embodiment of the present invention.

[0045] In this embodiment, the present invention can be applied to an application environment including, but not limited to, an electronic device 1 , a user terminal 2 , and a network 3 .

[0046] The electronic device 1 may be a computer, a server or other electronic device with data processing capability. In this embodiment, the electronic device is only an example of a server 1. In this embodiment, the server 1 may be a main server within an enterprise, or a server for managing a department in an enterprise. It may be a computing device such as a rack server, a blade server, a tower server or a cabinet server. The server 1 may be an independent server, or a server cluster composed of multiple servers.

[0047] The user terminal 2 in this embodiment may be a smart phone, a tablet computer, a personal computer, a portable computer, or other electronic devices with computing functions.

[0048] The network 3 may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), the Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, or Wi-Fi.

[0049] The server 1 is connected to the user terminal 2 through the network 3. The user terminal 2 receives the user image taken by the user and transmits it to the server 1 through the network 3. The server 1 processes the user image, specifically including: obtaining the user image to be processed, extracting the background image from the user image, extracting the first image feature from the background image, and compressing the first image feature to obtain the second image feature that can be used to perform image retrieval.

[0050] See also Figure 2 As shown, Figure 1 Schematic diagram of an optional hardware architecture of an electronic device 1. In this embodiment, the electronic device 1 may include, but is not limited to, a memory 11, a processor 12, and a network interface 13 that can be interconnected via a system bus. It should be noted that Figure 2 Only the electronic device 1 having components 11 - 13 is shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0051] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a hard disk or memory of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in hard disk equipped on the electronic device 1, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 11 can also include both the internal storage unit of the electronic device 1 and its external storage device. In this embodiment, the memory 11 is generally used to store the operating system and various application software installed on the electronic device 1, such as the program code of the image data processing device 100, etc. In addition, the memory 11 can also be used to temporarily store various types of data that have been output or are to be output.

[0052] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with the server 1. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the image data processing device 100.

[0053] The network interface 13 may include a wireless network interface or a wired network interface, and the network interface 13 is generally used to establish a communication connection between the electronic device 1 and other electronic devices. In this embodiment, the network interface 13 is mainly used to connect the electronic device 1 to the user terminal 2 through the network 3, and to establish a data transmission channel and a communication connection between the electronic device 1 and the user terminal 2.

[0054] See also Figure 3 , is a functional module diagram of an optional embodiment of the image data processing device 100 of the present invention. In this embodiment, the image data processing device 100 can be divided into one or more modules, one or more modules are stored in the memory 11, and are executed by one or more processors (processor 12 in this embodiment) to complete the present invention. For example, in Figure 3 In the embodiment, the image data processing device 100 can be divided into an image acquisition module 301, a background separation module 302, a feature extraction module 303 and a feature processing module 304. The functional module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions, which is more suitable for describing the execution process of the image data processing device 100 in the electronic device 1 than a program. The functions of each functional module 301-304 will be described in detail below.

[0055] The image acquisition module 301 is used to acquire a user image to be processed, wherein the user image includes a foreground portrait and a background image.

[0056] In this embodiment, for the convenience of description, a video interview in the field of financial fraud detection is used as an example for explanation. The user image acquired by the image acquisition module 301 is a user image submitted by a user who intends to start a certain business, and the business may be but is not limited to account opening, loan request, etc. The user image may be a real-time image extracted from a video interview in a financial business scenario, or may be uploaded after being shot by a user terminal.

[0057] The background separation module 302 is used to extract the background image from the user image.

[0058] In this embodiment, the background separation module 302 first uses a pre-trained segmentation model to detect the foreground portrait and background image from the user image. The pre-trained segmentation model described here can be, for example, SINet. SINet is a lightweight and robust portrait segmentation model. SINet includes an information block decoder and a spatial compression module. The information block decoder uses confidence estimation to re-cover local spatial information without destroying global consistency. The spatial compression module uses multiple receptive fields to process the consistency of various sizes in the image. The pre-trained segmentation model in this embodiment can also be other models, for example, it can also be PortraitNet, and the specific principle of PortraitNet will not be repeated.

[0059] After detecting the foreground portrait and the background image, the background separation module 302 sets the pixel value of the foreground portrait area to 0, sets the pixel value of the background image area to 1, and outputs a first two-dimensional matrix (e.g., a two-dimensional matrix mask) that is consistent with the size of the user image. Afterwards, the background separation module 302 multiplies the element value of the output first two-dimensional matrix mask with the pixel value of the corresponding pixel in the user image to obtain the background portrait with the foreground portrait subtracted. Figure 5 As an example, the background image with the foreground portrait subtracted is shown.

[0060] The feature extraction module 303 is used to extract the first image feature from the background image. In this embodiment, the feature extraction module is specifically used to: extract a deep feature map from the background image, divide the deep feature map into a preset number of feature map sub-blocks, and input the preset number of feature map sub-blocks into a point pair feature network (PPFNet) to obtain the first image feature.

[0061] In this embodiment, the feature extraction module 303 uses a lightweight network to extract a deep feature map from the background image.

[0062] Specifically, a lightweight network such as MobilenetV3 and shufflnetv2 is used to extract a H*W*C deep feature map from the background image, where H, W, and C are the height, width, and number of the deep feature map, respectively, and the specific parameters are related to the selected lightweight network. This embodiment uses a lightweight network to extract image features from the background image, which has the advantages of fewer network parameters, less involved calculations, and faster operation speed than using large networks such as resnet50 and vgg16 to extract image features. Moreover, the image features extracted by some large networks in the prior art are directly used as features for image retrieval, which cannot solve the problem of large-scale transformation, occlusion, and missing of user portraits and background information due to incomplete fixation of user terminals, such as mobile devices, in video face-to-face review scenarios. However, the image features extracted by the lightweight network in this embodiment are not ultimately used as features for image retrieval, and further processing is required.

[0063] In this embodiment, the feature extraction module 303 divides the deep feature map into a preset number of feature map sub-blocks.

[0064] Optionally, the deep feature map is divided into a grid of a preset size, such as an M*N grid, to obtain a preset number of feature map sub-blocks, namely, M*N feature map sub-blocks with a depth of C.

[0065] Optionally, the aspect ratio M / N of the *N grid is equal to the aspect ratio of the background image. For example, if the size of the background image is 320*240, M and N can be set to 4 and 3 (4 rows and 3 columns).

[0066] In this embodiment, the feature extraction module 303 further inputs a preset number of feature map sub-blocks into PPFNet to obtain the first image feature.

[0067] The existing PPFNet obtains global information 3D local feature descriptors through deep learning, thereby finding corresponding feature points in unordered point clouds. PPFNet has strong global context perception by learning local descriptors in pure geometric space. PPFNet uses point pair features, single points, and normal vectors calculated from neighborhoods as a whole to represent 3D information. At the same time, PPFNet uses a novel N-ary topological loss equation to naturally project global information onto local descriptors.

[0068] For the training of PPFNet, firstly, local small point cloud blocks are decomposed from the point cloud containing the common area, and the local blocks are input into PPFNet to obtain local features, and the feature similarity matrix of all features is calculated. The distance matrix between the local blocks is calculated through the real rigid posture of the two point clouds, and a correspondence matrix is ​​obtained by binarizing the distance matrix to identify all matching and non-matching relationships between the local blocks. Then, the N-ary topological loss equation is calculated by coupling the feature distance matrix and the correspondence matrix, so that PPFNet finds the optimal feature space.

[0069] In this embodiment, a preset number of M*N feature map sub-blocks are used as the input of PPFNet, and finally PPFNet outputs M*N one-dimensional local features based on global information with a preset dimension (for example, D dimension), that is, M*N*D dimensional features, as the first image feature. D represents the dimension of the feature output extracted by the network, and D can be set to 128 or 256. The conventional PPFNet data processing object is point cloud data, and the feature pyramid technology is used to process pixel-level point cloud data. Therefore, the amount of data calculation involved is huge, the calculation speed is slow, and the memory space is large. In this solution, the data object input into the PPFNet network is the feature map sub-block, and there is no need to process the pixel-level data, so problems such as high amount of calculation and large memory usage can be avoided.

[0070] The feature processing module 304 is used to compress the first image feature to obtain the final extracted second image feature. The second image feature can be used to perform image retrieval.

[0071] In this embodiment, the two-dimensional matrix mask is first scaled proportionally to the size of the feature map sub-block to obtain a new two-dimensional matrix mask', and the two-dimensional matrix mask' is multiplied by the first image feature to obtain a new M*N*D-dimensional feature, and then the M*N*D-dimensional feature is input into a pooling layer external to PPFNet and converted into a D-dimensional one-dimensional vector as the second image feature of the background image.

[0072] In this embodiment, the first image feature is compressed, and the above-mentioned M*N*D dimensional feature can also be input into a gem pooling layer, a maximum pooling layer, an average pooling layer, an rmac pooling layer, etc.

[0073] The second image features processed by the feature processing module 304 can be used to perform image retrieval. Performing image retrieval using the second image features can include, but is not limited to, further processing based on the second image features, such as clustering, similarity sorting, etc.

[0074] See also Figure 4 , is a flow chart of an embodiment of the image data processing method of the present invention. In this embodiment, according to different requirements, Figure 4The execution order of the steps in the flowchart shown may be changed, and some steps may be omitted.

[0075] Step S1: obtaining a user image to be processed, wherein the user image includes a foreground portrait and a background image.

[0076] In this step, for the sake of convenience, the video interview in the field of financial fraud detection is used as an example. The user image obtained in this step is the user image submitted by the user who intends to open a certain business, which can be but not limited to account opening, loan request, etc. The user image can be a real-time image extracted from the video interview of the financial business scene, or it can be a real-time image taken by the user terminal and uploaded.

[0077] Step S2: extracting a background image from the user image.

[0078] In this step, a pre-trained segmentation model is used to detect the foreground portrait and background image from the user image. The pre-trained segmentation model described here can be, for example, SINet. SINet is a lightweight and robust portrait segmentation model. SINet includes an information block decoder and a spatial compression module. The information block decoder uses confidence estimation to re-cover local spatial information without destroying global consistency. The spatial compression module uses multiple receptive fields to process the consistency of various sizes in the image. The pre-trained segmentation model in this embodiment can also be other models, for example, it can also be PortraitNet.

[0079] After detecting the foreground portrait and the background image, the pixel values ​​of the foreground portrait area are set to 0, the pixel values ​​of the background image area are set to 1, and a first two-dimensional matrix (e.g., a two-dimensional matrix mask) consistent with the size of the user image is output. Afterwards, the element values ​​of the output first two-dimensional matrix mask are multiplied by the pixel values ​​of the corresponding pixels in the user image to obtain the background portrait minus the foreground portrait.

[0080] Step S3, extracting a first image feature from the background image, comprising:

[0081] S31. Extract a deep feature map from the background image.

[0082] A lightweight network is used to extract deep feature maps from background images.

[0083] Specifically, a lightweight network such as MobilenetV3 and shufflnetv2 is used to extract a H*W*C deep feature map from the background image, where H, W, and C are the height, width, and number of the deep feature map, respectively, and the specific parameters are related to the selected lightweight network. This embodiment uses a lightweight network to extract image features from the background image, which has the advantages of fewer network parameters, less involved calculations, and faster operation speed than using large networks such as resnet50 and vgg16 to extract image features. Moreover, the image features extracted by some large networks in the prior art are directly used as features for image retrieval, which cannot solve the problem of large-scale transformation, occlusion, and missing of user portraits and background information due to incomplete fixation of user terminals, such as mobile devices, in video face-to-face review scenarios. However, the image features extracted by the lightweight network in this embodiment are not ultimately used as features for image retrieval, and further processing is required.

[0084] S32: Divide the deep feature map into a preset number of feature map sub-blocks.

[0085] In this step, the deep feature map is divided into a preset number of feature map sub-blocks. The deep feature map can be divided into a preset number of feature map sub-blocks by a grid of a preset size, such as an M*N grid, that is, M*N feature map sub-blocks with a depth of C.

[0086] Optionally, the aspect ratio M / N of the *N grid is equal to the aspect ratio of the background image. For example, if the size of the background image is 320*240, M and N can be set to 4 and 3 (4 rows and 3 columns).

[0087] S33. Input a preset number of feature map sub-blocks into PPFNet to obtain a first image feature.

[0088] In this step, a preset number of feature map sub-blocks are further input into PPFNet to obtain the first image feature. Specifically, a preset number of M*N feature map sub-blocks are used as the input of PPFNet, and finally PPFNet outputs M*N one-dimensional local features based on global information with a preset dimension (for example, D dimension), that is, M*N*D dimensional features, as the first image feature. D represents the dimension of the feature output extracted by the network, and D can be set to 128 or 256. The conventional PPFNet data processing object is point cloud data, and the feature pyramid technology is used to process pixel-level point cloud data. Therefore, the amount of data calculation involved is huge, the calculation speed is slow, and the memory space is large. In this solution, the data object input into the PPFNet network is the feature map sub-block, and there is no need to process the pixel-level data, so problems such as high amount of calculation and large memory usage can be avoided.

[0089] Step S4: compress the first image feature to obtain a second image feature that can be used to perform image retrieval.

[0090] In this step, the two-dimensional matrix mask is first scaled proportionally to the size of the feature map sub-block to obtain a new two-dimensional matrix mask', and the two-dimensional matrix mask' is multiplied by the first image feature to obtain a new M*N*D-dimensional feature, and then the M*N*D-dimensional feature is input into a gem pooling layer and converted into a D-dimensional one-dimensional vector as the second image feature of the background image.

[0091] In one embodiment of the present invention, the image data processing device stored in the memory 11 of the electronic device 1 is a combination of multiple computer programs, which, when executed in the processor 10, can achieve:

[0092] S1. Obtain a user image to be processed, wherein the user image includes a foreground portrait and a background image;

[0093] S2, extracting a background image from the user image;

[0094] S3, extracting a first image feature from the background image, specifically comprising:

[0095] Extracting a deep feature map from the background image;

[0096] Dividing the deep feature map into a preset number of feature map sub-blocks;

[0097] Inputting the preset number of feature map sub-blocks into PPFNet to obtain the first image features;

[0098] S4. Compress the first image feature to obtain a second image feature that can be used to perform image retrieval.

[0099] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium, and the computer-readable storage medium can be volatile or non-volatile. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0100] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0101] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0103] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0104] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0105] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The second and other words are used to indicate names, but not to indicate any particular order.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A method for processing image data, characterized in that: The method comprises: Acquire a user image to be processed, wherein the user image includes a foreground portrait and a background image; Extracting a background image from the user image; A deep feature map with H*W*C parameters is extracted from the background image using a lightweight network, the deep feature map is divided into M*N feature map sub-blocks, and the feature map sub-blocks are input into a point-to-point feature network, and a first image feature of M*N*D dimensions is output through the point-to-point feature network, where H, W, and C are the height, width, and number of the deep feature map, respectively; The first image feature is compressed to obtain a second image feature, and the second image feature is used to perform image retrieval.

2. The image data processing method according to claim 1, characterized in that: The extracting the background image from the user image comprises: Detecting a foreground person image and a background image from the user image using a pre-trained segmentation model; The pixel values ​​of the foreground portrait area are set to 0, the pixel values ​​of the background image area are set to 1, and a first two-dimensional matrix having the same size as the user image is output; The element values ​​of the first two-dimensional matrix are multiplied by the pixel values ​​of corresponding pixels in the user image to obtain the background image with the foreground portrait subtracted.

3. The image data processing method according to claim 1, characterized in that: The step of dividing the deep feature map into M*N feature map sub-blocks comprises: The deep feature map of the H*W*C parameters is evenly divided into M*N feature map sub-blocks with C layers through an M*N grid of a preset size.

4. The image data processing method according to claim 3, characterized in that: The aspect ratio M / N of the M*N grid is equal to the aspect ratio of the background image.

5. The image data processing method according to claim 1, characterized in that: The step of inputting the feature map sub-block into a point-to-point feature network and outputting a first image feature of M*N*D dimensions through the point-to-point feature network comprises: M*N feature map sub-blocks with C layers are used as input of the point-to-point feature network, and the point-to-point feature network outputs M*N one-dimensional local features based on global information with a preset D dimension, that is, M*N*D-dimensional features, as the first image features.

6. The image data processing method according to claim 2, characterized in that: The compressing the first image feature to obtain the second image feature includes: Scaling the first two-dimensional matrix in equal proportion to the size of the feature map sub-block to obtain a second two-dimensional matrix; Multiplying the second two-dimensional matrix by the first image feature to obtain a new M*N*D dimensional feature; The new M*N*D-dimensional feature is input into a pooling layer external to the point pair feature network and converted into a D-dimensional one-dimensional vector as the second image feature.

7. An image data processing device, characterized in that: The device comprises: An image acquisition module, used to acquire a user image to be processed, wherein the user image includes a foreground portrait and a background image; A background separation module, used to extract a background image from the user image; A feature extraction module, used for extracting a deep feature map with H*W*C parameters from the background image using a lightweight network, dividing the deep feature map into M*N feature map sub-blocks, and inputting the feature map sub-blocks into a point-to-point feature network, and outputting a first image feature of M*N*D dimensions through the point-to-point feature network, where H, W, and C are respectively the height, width, and number of the deep feature map; The feature processing module is used to compress the first image feature to obtain a second image feature, and use the second image feature to perform image retrieval.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores an image data processing program executable by the at least one processor, and the image data processing program is executed by the at least one processor so that the at least one processor can perform the image data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: An image data processing program is stored on the computer-readable storage medium, and the image data processing program can be executed by one or more processors to implement the image data processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Photo background similarity clustering method based on convolutional neural network and computer

    CN110569878A