Image processing method and device, electronic equipment and storage medium
By blocking and super-resolution reconstruction of continuous images, the problem of insufficient face detection accuracy in complex scenarios is solved, and higher image face detection accuracy is achieved.
Patent Information
- Application Number
- CN202510286700.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-18
Smart Images

Figure CN120340085A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] Face detection technology is a technology for identifying the position and size of a face in an image. In some complex scenarios, the face detection technology cannot accurately identify the face in the image.
[0003] In related technologies, for a to-be-detected image, single-frame super-resolution reconstruction processing is performed on the to-be-detected image to obtain super-resolution images of the to-be-detected image at multiple resolutions, and face detection is respectively performed on the super-resolution images at the above multiple resolutions to improve the accuracy of face detection for the to-be-detected image in complex scenarios.
[0004] However, the improvement effect of the accuracy of face detection for the to-be-detected image by the above method has limitations. Therefore, there is an urgent need to provide an image processing method to further improve the accuracy of face detection for the to-be-detected image in complex scenarios. Summary of the Invention
[0005] This application provides an image processing method, apparatus, electronic device, and storage medium for improving the accuracy of face detection for a to-be-detected image.
[0006] In a first aspect, this application provides an image processing method, which includes:
[0007] Obtain K consecutive frames of images, where K is a positive integer greater than 1;
[0008] Perform block processing on each of the K frames of images to obtain multiple first block images corresponding to each of the first K - 1 frames of images, and multiple second block images corresponding to the Kth frame of image;
[0009] For each second block image among the multiple second block images, perform super-resolution reconstruction processing on the second block image according to the multiple first block images corresponding to each of the first K - 1 frames of images to obtain multiple reconstructed images corresponding to the second block image, and the resolutions of the multiple reconstructed images are different;
[0010] Perform face detection on the Kth frame of image according to the multiple reconstructed images corresponding to each of the multiple second block images to obtain the face detection result of the Kth frame of image.
[0011] In a possible implementation manner, performing super-resolution reconstruction processing on the second block image according to the multiple first block images corresponding to each of the first K - 1 frames of images to obtain multiple reconstructed images corresponding to the second block image includes:
[0012] Determine the super-resolution reconstruction method of the second sub-block image according to multiple first sub-block images respectively corresponding to the first K-1 frame images; the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method, or a single-frame super-resolution reconstruction method;
[0013] Perform super-resolution reconstruction processing on the second sub-block image according to the super-resolution reconstruction method of the second sub-block image to obtain multiple reconstructed images corresponding to the second sub-block image.
[0014] In a possible implementation manner, determining the super-resolution reconstruction method of the second sub-block image according to multiple first sub-block images respectively corresponding to the first K-1 frame images includes:
[0015] Determine target pixel points in the second sub-block image;
[0016] For each of the first K-1 frame images, perform matching processing on the multiple first sub-block images corresponding to the image and the second sub-block image, determine the target sub-block image in the multiple first sub-block images corresponding to the image, and the position of the target pixel point in the target sub-block image, where the target sub-block image is the first sub-block image to which the target pixel point belongs among the multiple first sub-block images corresponding to the image;
[0017] Determine the movement trajectory of the target pixel point according to the target sub-block images respectively corresponding to the first K-1 frame images and the positions of the target pixel points in the target sub-block images respectively corresponding to the first K-1 frame images;
[0018] Determine the super-resolution reconstruction method of the second sub-block image according to the movement trajectory of the target pixel point.
[0019] In a possible implementation manner, determining the super-resolution reconstruction method of the second sub-block image according to the movement trajectory of the target pixel point includes:
[0020] Determine the movement amplitude between each image in the first K-1 frame images and the Kth frame image according to the movement trajectory of the target pixel point;
[0021] Determine the target movement amplitude according to the movement amplitudes between each image in the first K-1 frame images and the Kth frame image; the target movement amplitude is the average value of the movement amplitudes between each image in the first K-1 frame images and the Kth frame image, or the maximum value of the movement amplitudes between each image in the first K-1 frame images and the Kth frame image;
[0022] When the target movement amplitude is less than or equal to the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is a multi-frame super-resolution reconstruction method;
[0023] When the target movement amplitude is greater than the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is a single-frame super-resolution reconstruction method.
[0024] In a possible implementation, the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method. The second sub-block image is subjected to super-resolution reconstruction processing to obtain multiple reconstructed images corresponding to the second sub-block image, including:
[0025] The second sub-block image and the target sub-block images corresponding to the previous K - 1 frames of images are subjected to fusion processing to obtain a fused image corresponding to the second sub-block image;
[0026] According to multiple preset magnification factors, the fused image is subjected to super-resolution reconstruction processing to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0027] In a possible implementation, the super-resolution reconstruction method is a single-frame super-resolution reconstruction method. The second sub-block image is subjected to super-resolution reconstruction processing to obtain multiple reconstructed images corresponding to the second sub-block image, including:
[0028] According to multiple preset magnification factors, the second sub-block image is subjected to super-resolution reconstruction to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0029] In a possible implementation, based on the multiple reconstructed images corresponding to each of the multiple second sub-block images, face detection is performed on the Kth frame image to obtain the face detection result of the Kth frame image, including:
[0030] Face detection processing is performed on the Kth frame image to determine the first detection box in the Kth frame image;
[0031] For each second sub-block image, according to the multiple reconstructed images of the second sub-block image, a second detection box corresponding to the second sub-block image is determined in the Kth frame image;
[0032] According to the position of the first detection box in the Kth frame image, the confidence of the first detection box, the positions of the second detection boxes corresponding to each second sub-block image, and the confidence of the second detection boxes, screening processing is performed on the first detection box and the second detection boxes corresponding to each second sub-block image to obtain the target detection box in the Kth frame image;
[0033] Wherein, the face detection result includes the position and size of the target detection box.
[0034] In a possible implementation, determining the second detection box corresponding to the second sub-block image in the Kth frame image according to the multiple reconstructed images of the second sub-block image includes:
[0035] For each of the multiple reconstructed images, face detection processing is performed on the reconstructed image to obtain a third detection box in the reconstructed image;
[0036] Determine a second detection box in the K-th frame image according to the magnification of each reconstructed image relative to the second sub-block image, the positions of each third detection box in the reconstructed image, and the position of the second sub-block image in the K-th frame image.
[0037] In a second aspect, the present application provides an image processing apparatus, including:
[0038] An acquisition module, configured to acquire K consecutive frame images, where K is a positive integer greater than 1;
[0039] A sub-block module, configured to perform sub-block processing on the K frame images respectively, to obtain a plurality of first sub-block images corresponding to each of the first K - 1 frame images, and a plurality of second sub-block images corresponding to the K-th frame image;
[0040] A reconstruction module, configured to perform super-resolution reconstruction processing on each second sub-block image among the plurality of second sub-block images according to the plurality of first sub-block images corresponding to each of the first K - 1 frame images, to obtain a plurality of reconstructed images corresponding to the second sub-block image, and the resolutions of the plurality of reconstructed images are different;
[0041] A detection module, configured to perform face detection on the K-th frame image according to the plurality of reconstructed images corresponding to each of the plurality of second sub-block images, to obtain the face detection result of the K-th frame image.
[0042] In a possible implementation manner, the reconstruction module is specifically configured to:
[0043] Determine the super-resolution reconstruction method of the second sub-block image according to the plurality of first sub-block images corresponding to each of the first K - 1 frame images; the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method, or a single-frame super-resolution reconstruction method;
[0044] Perform super-resolution reconstruction processing on the second sub-block image according to the super-resolution reconstruction method of the second sub-block image, to obtain a plurality of reconstructed images corresponding to the second sub-block image.
[0045] In a possible implementation manner, the reconstruction module is specifically configured to:
[0046] Determine target pixel points in the second sub-block image;
[0047] For each of the first K - 1 frame images, perform matching processing on the plurality of first sub-block images corresponding to the image and the second sub-block image, determine a target sub-block image in the plurality of first sub-block images corresponding to the image, and the position of the target pixel points in the target sub-block image, where the target sub-block image is the first sub-block image to which the target pixel points belong among the plurality of first sub-block images corresponding to the image;
[0048] Determine the movement trajectory of the target pixel point according to the target sub-block images corresponding to the respective first K-1 frame images and the positions of the target pixel point in the target sub-block images corresponding to the respective first K-1 frame images;
[0049] Determine the super-resolution reconstruction method of the second sub-block image according to the movement trajectory of the target pixel point.
[0050] In a possible implementation manner, the reconstruction module is specifically configured to:
[0051] Determine the movement amplitude between each of the first K-1 frame images and the Kth frame image according to the movement trajectory of the target pixel point;
[0052] Determine the target movement amplitude according to the movement amplitudes between each of the first K-1 frame images and the Kth frame image; the target movement amplitude is the average value of the movement amplitudes between each of the first K-1 frame images and the Kth frame image, or the maximum value of the movement amplitudes between each of the first K-1 frame images and the Kth frame image;
[0053] In the case where the target movement amplitude is less than or equal to the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the multi-frame super-resolution reconstruction method;
[0054] In the case where the target movement amplitude is greater than the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the single-frame super-resolution reconstruction method.
[0055] In a possible implementation manner, the reconstruction module is specifically configured to:
[0056] Perform a fusion process on the second sub-block image and the target sub-block images corresponding to the respective first K-1 frame images to obtain a fused image corresponding to the second sub-block image;
[0057] Perform super-resolution reconstruction processing on the fused image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0058] In a possible implementation manner, the reconstruction module is specifically configured to:
[0059] Perform super-resolution reconstruction on the second sub-block image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0060] In a possible implementation manner, the detection module is specifically configured to:
[0061] Perform face detection processing on the Kth frame image to determine the first detection frame in the Kth frame image;
[0062] For each second sub-block image, according to multiple reconstructed images of the second sub-block image, determine a second detection box corresponding to the second sub-block image in the K-th frame image;
[0063] According to the position of the first detection box in the K-th frame image, the confidence of the first detection box, the positions of the second detection boxes corresponding to the second sub-block images, and the confidence of the second detection boxes, perform screening processing on the first detection box and the second detection boxes corresponding to the second sub-block images to obtain a target detection box in the K-th frame image;
[0064] Among them, the face detection result includes the position and size of the target detection box.
[0065] In a possible implementation manner, the detection module is specifically configured to:
[0066] For each reconstructed image among the multiple reconstructed images, perform face detection processing on the reconstructed image to obtain a third detection box in the reconstructed image;
[0067] According to the magnification of each reconstructed image relative to the second sub-block image, the position of each third detection box in the reconstructed image, and the position of the second sub-block image in the K-th frame image, determine the second detection box in the K-th frame image.
[0068] In a third aspect, the present application provides an electronic device, including:
[0069] At least one processor; and
[0070] A memory communicatively connected to the at least one processor; wherein,
[0071] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image processing method involved in the first aspect and any possible implementation manner.
[0072] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, wherein computer-executable instructions are stored in the computer-readable storage medium, and the computer-executable instructions are used to cause a computer to execute the image processing method involved in the first aspect and any possible implementation manner.
[0073] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the image processing method involved in the first aspect and any possible implementation manner.
[0074] In a sixth aspect, the present application provides a chip, the chip includes at least one processor, and the processor is used to run program instructions to execute the image processing method involved in the first aspect and any possible implementation manner.
[0075] The image processing method, apparatus, electronic device, and storage medium provided by the embodiments of this application acquire K consecutive frames of images, block the first K - 1 frames of images to obtain a plurality of first block images; block the Kth frame of image to obtain a plurality of second block images. For each second block image, according to the first block images corresponding to the first K - 1 frames of images respectively, perform super - resolution reconstruction processing on the second block image to obtain reconstructed images of the second block image at multiple resolutions. Then, according to the reconstructed images of the second block image at multiple resolutions, perform face detection on the Kth frame of image to obtain the face detection result of the Kth frame of image. In the above method, performing super - resolution reconstruction processing on the second block image according to the second block image and the first block images corresponding to the first K - 1 frames of images respectively can effectively utilize the redundant information of the second block image in time and space, improve the resolution and detail performance of the second block image, so that the obtained reconstructed images have higher resolution and stronger information expression ability. Then, when performing face detection on the Kth frame of image according to the reconstructed images of the second block image at multiple resolutions, the obtained face detection result will be more accurate, thereby improving the accuracy of face detection for the Kth frame of image. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application and used together with the specification to explain the principles of this application.
[0077] Figure 1 It is a schematic diagram of the application scenario provided by the embodiments of this application;
[0078] Figure 2 It is a schematic flowchart of an image processing method provided by the embodiments of this application;
[0079] Figure 3 It is a schematic flowchart of performing super - resolution reconstruction processing on a second block image provided by the embodiments of this application;
[0080] Figure 4 It is a schematic flowchart of performing face detection on the Kth frame of image provided by the embodiments of this application;
[0081] Figure 5 It is a schematic diagram of the structure of an image processing apparatus provided by the embodiments of this application;
[0082] Figure 6 It is a schematic diagram of the structure of the electronic device provided by the embodiments of this application.
[0083] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0084] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0085] In the technical solution of the present application, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of user data and other information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0086] It should be noted that in the embodiments of the present application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0087] For ease of understanding, the following combines Figure 1 to briefly describe the application scenarios applicable to the embodiments of the present application.
[0088] Figure 1 FIG. is a schematic diagram of the application scenario provided for the embodiments of the present application. Please refer to Figure 1 , which includes an Internet of Things device 11 and a server 12. The server 12 is connected to the Internet of Things device 11 through a wired or wireless connection. The Internet of Things device 11 can be, for example, a device with a camera function such as a mobile phone camera, a surveillance camera, a camera, a vehicle-mounted camera, or an access control camera.
[0089] The Internet of Things device 11 can send a to-be-detected image 13 to the server 12. After receiving the to-be-detected image 13, the server 12 performs image processing on the to-be-detected image 13.
[0090] It should be noted that the Internet of Things device 11 and the server 12 can be two independent devices, or two different components integrated in the same device. In the Figure 1 example, the Internet of Things device 11 and the server 12 are taken as two independent devices for introduction.
[0091] Although in Figure 1In the example, the execution entity is server 12. However, the execution entity in the embodiments of this application can be a device with data processing functions such as a processor, a server, a microprocessor, a chip, etc. For example, the execution entity can also be the Internet of Things device 11. The specific execution entity in each embodiment of this application is not limited, and it can be selected and set according to actual needs. As long as it is a device with data processing functions, it can be used as the execution entity in each embodiment of this application.
[0092] It should be noted that Figure 1 it is only an example to illustrate an application scenario, and it is not a limitation on the application scenario.
[0093] Face detection technology is a technology for determining the position and size of a face in an image. For face detection in some complex scenarios, the accuracy of face detection is relatively low. Face detection in complex scenarios can be, for example, small-scale face detection, large-pose face detection, low-light face detection, blurred face detection, etc. Taking small-scale face detection as an example, the size of the face in the image to be detected accounts for a relatively small proportion. In this way, the face information will be compressed. When performing face detection on the image to be detected, due to the inability to obtain sufficient face information, problems such as missed detection may occur.
[0094] In the related art, for face detection in complex scenarios, multiple resolutions corresponding to the image to be detected are determined, and single-frame super-resolution reconstruction is performed on the image to be detected according to the multiple resolutions to obtain super-resolution images of the image to be detected at multiple resolutions. Then, face detection processing is performed on the super-resolution images of the image to be detected at multiple resolutions to obtain a face detection image of the image to be detected, thereby improving the accuracy of face detection of the image to be detected. However, since the face information included in the image to be detected is limited, the effect of performing super-resolution reconstruction on the image to be detected to obtain super-resolution images of the image to be detected at multiple resolutions is poor. Therefore, there are limitations in the effect of improving face detection accuracy by performing face detection on the super-resolution images of the image to be detected at multiple resolutions.
[0095] The face detection method provided by the embodiments of this application acquires K consecutive frames of images, divides the first K - 1 frames of images into blocks to obtain a plurality of first block images; divides the Kth frame of image into blocks to obtain a plurality of second block images. For each second block image, according to the first block images corresponding to the first K - 1 frames of images respectively, super-resolution reconstruction processing is performed on the second block image, which can effectively utilize the redundant information of the second block image in time and space, improve the resolution and detail performance of the second block image, and make the obtained reconstructed image have a higher resolution and stronger information expression ability. Then, according to the reconstructed images of the second block image at multiple resolutions, face detection is performed on the Kth frame of image, thereby improving the accuracy of face detection for the Kth frame of image.
[0096] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of this application with reference to the accompanying drawings.
[0097] Figure 2 It is a schematic flowchart of an image processing method provided by the embodiments of this application. Please refer to Figure 2 , this method may include the following steps:
[0098] S21. Acquire K consecutive frames of images, where K is a positive integer greater than 1.
[0099] In some scenarios, it is necessary to perform face detection processing on the image to be detected to determine the size and position of the face in the image to be detected. The K frames of images include the image to be detected, and the image to be detected is the last frame of the K frames of images. In the K frames of images, the size of each frame of image is the same. Exemplarily, assume that K is 5, and the K frames of images are the 1st frame image, the 2nd frame image, the 3rd frame image, the 4th frame image, and the 5th frame image in sequence. Then the image to be detected is the 5th frame image, and the sizes of the 1st frame image, the 2nd frame image, the 3rd frame image, the 4th frame image, and the 5th frame image are all the same.
[0100] It should be noted that in the case where the number of frames of the existing multi-frame images is less than K, the 1st frame image in the continuous multi-frame images is used to supplement the continuous multi-frame images so that the number of frames of the multi-frame images is equal to K. Exemplarily, assume that 3 consecutive frames of images are acquired, and the 3 consecutive frames of images are the 1st frame image, the 2nd frame image, and the 3rd frame image in sequence, and K is 5. The number of frames of the acquired continuous multi-frame images is less than K. At this time, the 1st frame image is used to supplement the 3 consecutive frames of images, and the 5 frames of images obtained after supplementation are the 1st frame image, the 1st frame image, the 1st frame image, the 2nd frame image, and the 3rd frame image in sequence.
[0101] S22. Perform block processing on each of the K-frame images to obtain multiple first block images corresponding to each of the first K-1 frame images and multiple second block images corresponding to the K-th frame image.
[0102] The first K-1 frame images refer to the K-1 frame images among the consecutive K-frame images excluding the K-th frame image. Exemplarily, assuming K is 3, and the consecutive K-frame images are the 1st frame image, the 2nd frame image, and the 3rd frame image in sequence, then the first 2 frame images are the 1st frame image and the 2nd frame image respectively.
[0103] The first block image is the block image obtained after performing block processing on the first K-1 frame images; the second block image is the block image obtained after performing block processing on the K-th frame image. In some embodiments, the first block image and the second block image can be rectangular or square. Taking the case where both the first block image and the second block image are rectangular as an example, the size of the first block image can be represented by the length and width of the first block image; the size of the second block image can be represented by the length and width of the second block image, and the size of the first block image is the same as that of the second block image.
[0104] For any one of the K-frame images, performing block processing on this image means dividing this image according to the size of the block image, and the number of block images obtained after division is the number of blocks of this image in the row direction multiplied by the number of blocks of this image in the column direction. Since the size of each frame image is the same and the size of each block image is the same, therefore, after performing block processing on each frame image, the number of blocks corresponding to each frame image is consistent.
[0105] It should be noted that during the process of dividing any one of the K-frame images into blocks, a certain coverage area needs to be reserved at the boundary of the block image, so that a certain scale of face information can be ensured to be included in the block image.
[0106] Exemplarily, assuming K is 3, the K-frame images are the 1st frame image, the 2nd frame image, and the 3rd frame image respectively, and the sizes of the 1st frame image, the 2nd frame image, and the 3rd frame image are 20*20, and the size of the block image is 5*5. Taking the division of the 1st frame image as an example, the number of blocks of the 1st frame image in the row direction is 4, and the number of blocks of the 1st frame image in the column direction is 4, then the number of first block images obtained after dividing the 1st frame image is 4*4; taking the division of the 3rd frame image as an example, the number of blocks of the 3rd frame image in the row direction is 4, and the number of blocks of the 3rd frame image in the column direction is 4, then the number of second block images obtained after dividing the 3rd frame image is 4*4.
[0107] S23. For each of the multiple second sub-block images, perform super-resolution reconstruction processing on the second sub-block image according to the multiple first sub-block images corresponding to each of the previous K-1 frame images, to obtain multiple reconstructed images corresponding to the second sub-block image, where the resolutions of the multiple reconstructed images are different.
[0108] Performing super-resolution reconstruction processing on the second sub-block image means performing multiple reconstructions on the second sub-block image according to different resolutions to obtain reconstructed images of the second sub-block image at different resolutions. Exemplarily, assume the resolution of the second sub-block image is 100*100. Perform super-resolution reconstruction processing on the second sub-block image according to the resolution of 200*200 to obtain a reconstructed image of the second sub-block image at the resolution of 200*200; perform super-resolution reconstruction processing on the second sub-block image according to the resolution of 300*300 to obtain a reconstructed image of the second sub-block image at the resolution of 300*300.
[0109] Performing super-resolution reconstruction processing on the second sub-block image according to the multiple first sub-block images corresponding to each of the previous K-1 frame images can effectively utilize the redundant information in time and space, improve the resolution and detail performance of the second sub-block image. Therefore, after performing super-resolution reconstruction on the second sub-block image, the obtained reconstructed images have higher resolutions and stronger information expression capabilities.
[0110] S24. Perform face detection on the Kth frame image according to the multiple reconstructed images corresponding to each of the multiple second sub-block images, to obtain the face detection result of the Kth frame image.
[0111] Performing face detection on the Kth frame image according to the multiple reconstructed images corresponding to each of the multiple second sub-block images means performing face detection on the multiple reconstructed images corresponding to each of the multiple second sub-block images respectively to obtain preliminary face detection results on the multiple reconstructed images, and then fusing and screening the preliminary face detection results on the multiple reconstructed images, so as to obtain the face detection result of the Kth frame image.
[0112] At Figure 2In the illustrated embodiment, the first K-1 frame images are partitioned to obtain a plurality of first partitioned images; the K-th frame image is partitioned to obtain a plurality of second partitioned images. For each second partitioned image, super-resolution reconstruction processing is performed on the second partitioned image according to the first partitioned images corresponding to the first K-1 frame images respectively, to obtain reconstructed images of the second partitioned image at multiple resolutions. Then, based on the reconstructed images of the second partitioned image at multiple resolutions, face detection is performed on the K-th frame image to obtain the face detection result of the K-th frame image. In the above method, by performing super-resolution reconstruction processing on the second partitioned image according to the second partitioned image and the first partitioned images corresponding to the first K-1 frame images respectively, the redundant information of the second partitioned image in time and space can be effectively utilized, the resolution and detail performance of the second partitioned image can be improved, and the obtained reconstructed images have higher resolution and stronger information expression ability. Then, when face detection is performed on the K-th frame image based on the reconstructed images of the second partitioned image at multiple resolutions, the obtained face detection result will be more accurate, thereby improving the accuracy of face detection for the K-th frame image.
[0113] In Figure 2 the illustrated embodiment, it is introduced that before performing face detection on the K-th frame image, super-resolution reconstruction processing is performed on the second partitioned image according to the plurality of first partitioned images corresponding to the first K-1 frame images respectively. Next, in combination with Figure 3 , the process of performing super-resolution reconstruction processing on the second partitioned image is further described.
[0114] Figure 3 FIG. is a schematic flow chart of performing super-resolution reconstruction processing on a second partitioned image provided by an embodiment of the present application. Please refer to Figure 3 , the method may include the following steps:
[0115] S31. Determine the super-resolution reconstruction method of the second partitioned image according to the plurality of first partitioned images corresponding to the first K-1 frame images respectively; the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method or a single-frame super-resolution reconstruction method.
[0116] The multi-frame super-resolution reconstruction method means that super-resolution reconstruction processing is performed on the second partitioned image according to the plurality of first partitioned images corresponding to the first K-1 frame images respectively.
[0117] The single-frame super-resolution reconstruction method means that super-resolution reconstruction processing is directly performed on the second partitioned image.
[0118] The method for determining the super-resolution reconstruction of the second sub-block image can be as follows: Determine the target pixel points in the second sub-block image; for each of the first K-1 frames of images, perform matching processing on the corresponding multiple first sub-block images and the second sub-block image, determine the target sub-block image in the corresponding multiple first sub-block images of the image, and the position of the target pixel point in the target sub-block image. The target sub-block image is the first sub-block image to which the target pixel point belongs among the corresponding multiple first sub-block images of the image; Determine the movement trajectory of the target pixel point according to the target sub-block images corresponding to the first K-1 frames of images respectively, and the positions of the target pixel point in the target sub-block images corresponding to the first K-1 frames of images respectively; Determine the super-resolution reconstruction method of the second sub-block image according to the movement trajectory of the target pixel point.
[0119] The target pixel points are the pixel points in the second sub-block image. In some embodiments, for any second sub-block image, the target pixel points can be randomly determined in the second sub-block image, or the pixel points at the specified positions in the second sub-block image can be determined as the target pixel points. Alternatively, by methods such as feature point detection, the pixel points with significant features can be detected in the second sub-block image, and then the pixel points with significant features can be determined as the target pixel points.
[0120] In some embodiments, for each of the first K-1 frames of images, the method for performing matching processing on the corresponding multiple first sub-block images and the second sub-block image can be methods such as feature matching method, optical flow method, block matching, and phase correlation method.
[0121] Taking the feature matching method as an example, for each of the first K-1 frames of images, the method for performing matching processing on the corresponding multiple first sub-block images and the second sub-block image, and determining the target sub-block image in the corresponding multiple first sub-block images of the image, and the position of the target pixel point in the target sub-block image can be as follows: For any second sub-block image, determine the target sub-block image in the multiple first sub-block images according to the position of the second sub-block image in the Kth frame of image. Then, for any target pixel point in the second sub-block image, extract the local features around the target pixel point in the second sub-block image. In the target sub-block image, find the features similar to the local features around the target pixel point, and then determine the position of the target pixel point in the target sub-block image.
[0122] Exemplarily, assume that K is 3, and the consecutive K-frame images are the first-frame image, the second-frame image, and the third-frame image in sequence. The multiple first sub-block images corresponding to the first-frame image are the first sub-block image 1 and the first sub-block image 2, where the position of the first sub-block image 1 in the first-frame image is the left-side sub-block image, and the position of the first sub-block image 2 in the first-frame image is the right-side sub-block image; the multiple first sub-block images corresponding to the second-frame image are the first sub-block image 3 and the first sub-block image 4, where the position of the first sub-block image 3 in the second-frame image is the left-side sub-block image, and the position of the first sub-block image 4 in the second-frame image is the right-side sub-block image; the multiple second sub-block images corresponding to the third-frame image are the second sub-block image 5 and the second sub-block image 6, where the position of the second sub-block image 5 in the third-frame image is the left-side sub-block image, and the position of the second sub-block image 6 in the third-frame image is the right-side sub-block image. Taking the second sub-block image 5 as an example, since the positions of the first sub-block image 1 in the first-frame image, the first sub-block image 3 in the second-frame image, and the second sub-block image 5 in the third-frame image are the same, the target sub-block images of the second sub-block image 5 are determined to be the first sub-block image 1 and the first sub-block image 3. Determine the target pixel point A in the second sub-block image 5, and extract the local features around the pixel point A in the second sub-block image 5. Then, respectively find the features similar to the local features around the pixel point A in the first sub-block image 1 and the first sub-block image 3. According to the position of the feature similar to the local features around the pixel point A in the first sub-block image 1, determine the position of the target pixel point A in the first sub-block image 1; according to the position of the feature similar to the local features around the pixel point A in the first sub-block image 3, determine the position of the target pixel point A in the first sub-block image 3.
[0123] After determining the target sub-block images and the positions of the target pixel points in the target sub-block images, according to the positions of the target pixel points in the target sub-block images, determine the movement trajectories of the target pixel points. The methods for determining the movement trajectories of the target pixel points can be methods such as linear trajectory fitting, polynomial trajectory fitting, and particle filtering. Then, according to the movement trajectories of the target pixel points, determine the super-resolution reconstruction method of the second sub-block image.
[0124] In some embodiments, the multiple first sub-block images and the second sub-block images corresponding to the image can be matched, and the movement trajectories between the target pixel points can be determined through a trained neural network model.
[0125] The method for determining the super-resolution reconstruction method of the second sub-block image according to the moving trajectory of the target pixel can be as follows: According to the moving trajectory of the target pixel, determine the moving amplitude between each of the first K-1 frames of images and the Kth frame of image; According to the moving amplitudes between each of the first K-1 frames of images and the Kth frame of image, determine the target moving amplitude; The target moving amplitude is the average value of the moving amplitudes between each of the first K-1 frames of images and the Kth frame of image, or the maximum value of the moving amplitudes between each of the first K-1 frames of images and the Kth frame of image; In the case where the target moving amplitude is less than or equal to the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the multi-frame super-resolution reconstruction method; In the case where the target moving amplitude is greater than the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the single-frame super-resolution reconstruction method.
[0126] For any one of the first K-1 frames of images, the moving amplitude between this image and the Kth frame of image is used to represent the motion intensity between this image and the Kth frame of image. The greater the moving amplitude, the greater the motion intensity between this image and the Kth frame of image; The smaller the moving amplitude, the smaller the motion intensity between this image and the Kth frame of image. In some embodiments, the moving amplitude between this image and the Kth frame of image can be represented by the angle and distance of the target pixel between this image and the Kth frame of image. Therefore, according to the moving trajectory of the target pixel between this image and the Kth frame of image, the moving amplitude of the target pixel between this image and the Kth frame of image can be determined, and then the moving amplitude between this image and the Kth frame of image can be determined.
[0127] Exemplarily, assume that the first K-1 frames of images are the first frame of image, the second frame of image respectively, and the Kth frame of image is the third frame of image. Taking the determination of the moving amplitude between the first frame of image and the third frame of image as an example for illustration. There is a target pixel A. According to the moving trajectory of the target pixel A between the first frame of image and the third frame of image, the moving amplitude of the target pixel A between the first frame of image and the third frame of image is determined as: the moving angle is 90 degrees clockwise, and the moving distance is 0.1 meters. Then the moving amplitude between the first frame of image and the third frame of image is determined as: the moving angle is 90 degrees clockwise, and the moving distance is 0.1 meters.
[0128] In the case where there are multiple target pixels, for any one of the first K-1 frames of images, the moving amplitudes of the multiple target pixels between this image and the Kth frame of image can be determined according to the moving trajectories of the multiple target pixels between this image and the Kth frame of image respectively. Then, by calculating the average value of the moving amplitudes of the multiple target pixels between this image and the Kth frame of image respectively, or by taking the maximum value of the moving amplitudes of the multiple target pixels between this image and the Kth frame of image and other methods, the moving amplitude between two adjacent images is determined.
[0129] Exemplarily, assume that the first K-1 frame images are the first frame image, the second frame image, and the Kth frame image is the third frame image. Taking the determination of the movement amplitude between the first frame image and the third frame image as an example for illustration. There are target pixel points A and B. The movement amplitude of target pixel point A between the first frame image and the third frame image can be determined according to the movement trajectory of target pixel point A between the first frame image and the third frame image: the movement angle is 90 degrees clockwise, and the movement distance is 0.1 meters; according to the movement trajectory of target pixel point B between the first frame image and the third frame image, the movement amplitude of target pixel point B between the first frame image and the third frame image is determined as: the movement angle is 50 degrees clockwise, and the movement distance is 0.1 meters. Calculate the average value of the movement amplitude of target pixel point A between the first frame image and the third frame image and the movement amplitude of target pixel point B between the first frame image and the third frame image, and determine the movement amplitude between the first frame image and the third frame image as: the movement angle is 50 degrees clockwise, and the movement distance is 0.1 meters.
[0130] Then, according to the movement amplitudes between each of the first K frame images and the Kth frame image, determine the target movement amplitude, which is used to represent the overall motion intensity of the K frame images.
[0131] The smaller the target movement amplitude, the smoother the overall motion of the K frame images. When the overall motion of the K frame images is relatively smooth, the correlation between the first K-1 frame images and the Kth frame image is relatively large. At this time, performing super-resolution reconstruction processing on the second sub-block image according to the multiple first sub-block images corresponding to the first K-1 frame images respectively, the obtained reconstructed image has a better effect. Therefore, when the target movement amplitude is less than or equal to the preset amplitude, determine the super-resolution reconstruction method of the second sub-block image as the multi-frame super-resolution reconstruction method.
[0132] The larger the target movement amplitude, the stronger the overall motion of the K frame images. When the overall motion of the K frame images is relatively strong, the correlation between the first K-1 frame images and the Kth frame image is relatively small. At this time, performing super-resolution reconstruction processing on the second sub-block image according to the multiple first sub-block images corresponding to the first K-1 frame images respectively, the obtained reconstructed image has a poor effect. Therefore, when the target movement amplitude is greater than the preset amplitude, determine the super-resolution reconstruction method of the second sub-block image as the single-frame super-resolution reconstruction method.
[0133] S32. Perform super-resolution reconstruction processing on the second sub-block image according to the super-resolution reconstruction method of the second sub-block image to obtain multiple reconstructed images corresponding to the second sub-block image.
[0134] When the super-resolution reconstruction method is the multi-frame super-resolution reconstruction method, the method for performing super-resolution reconstruction processing on the second sub-block image can be as follows: perform fusion processing on the second sub-block image and the target sub-block images corresponding to the previous K-1 frames of images respectively to obtain a fused image corresponding to the second sub-block image; perform super-resolution reconstruction processing on the fused image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0135] In some embodiments, the method for performing fusion processing on the second sub-block image and the target sub-block images corresponding to the previous K-1 frames of images respectively can be as follows: the second sub-block image and the target sub-block images corresponding to the previous K-1 frames of images can be spatially aligned by methods such as affine transformation and perspective transformation. Then, for any pixel point in the second sub-block image, determine the pixel value of this pixel point in the target sub-block images corresponding to the previous K-1 frames of images respectively, and calculate the weighted average value of the pixel value of this pixel point in the second sub-block image and the pixel values of this pixel point in each target sub-block image, so as to achieve the fusion of the second sub-block image and the target sub-block images corresponding to the previous K-1 frames of images respectively. Finally, perform denoising and other processing on the fused image to obtain a fused image corresponding to the second sub-block image.
[0136] The resolutions of the multiple reconstructed images can be determined according to the preset magnification factors, and then, according to the resolutions of the multiple reconstructed images respectively, perform super-resolution reconstruction processing on the fused image to obtain reconstructed images of the second sub-block image at multiple resolutions. The more the preset magnification factors are, the more the reconstructed images corresponding to each sub-block image are. Therefore, in some embodiments, if it is necessary to save computing resources, energy consumption, etc., the number of preset magnification factors can be appropriately reduced.
[0137] Exemplarily, assume that the preset magnification factors are 2 and 3, and the resolution of the fused image is 20*20. Then, determine the resolutions of the multiple reconstructed images to be 40*40 and 60*60 respectively. Perform super-resolution reconstruction on the fused image according to the resolution 40*40 to obtain reconstructed image 1; perform super-resolution reconstruction on the fused image according to the resolution 60*60 to obtain reconstructed image 2.
[0138] In some embodiments, super-resolution reconstruction of the fused image can be performed by methods such as interpolation-based methods and reconstruction-based methods. Among them, interpolation-based methods can be, for example, nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Taking the bilinear interpolation method as an example, the method for performing super-resolution reconstruction on the fused image using the bilinear interpolation method can be as follows: Determine the position of the pixel point to be interpolated in the fused image, and determine the pixel values of the 4 pixel points around the pixel point. Calculate the weighted average of the pixel values of the 4 pixel points around the pixel point respectively, and then determine the weighted average as the pixel value of the pixel point. Repeat the above operations until the resolution of the reconstructed image is the resolution determined according to the preset magnification.
[0139] In some embodiments, multi-frame super-resolution reconstruction processing can be performed on the second sub-block image through a trained multi-frame super-resolution reconstruction model.
[0140] When the super-resolution reconstruction method is a single-frame super-resolution reconstruction method, the method for performing super-resolution reconstruction processing on the second sub-block image can be as follows: Perform super-resolution reconstruction on the second sub-block image according to multiple preset magnifications to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0141] In some embodiments, super-resolution reconstruction of the second sub-block image can be performed by methods such as interpolation-based methods and reconstruction-based methods. Taking the bicubic interpolation method as an example, the method for performing super-resolution reconstruction on the second sub-block image using the bicubic interpolation method can be as follows: Determine the position of the pixel point to be interpolated in the second sub-block image, and determine the pixel values of the 16 pixel points around the pixel point. Calculate the weighted average of the pixel values of the 16 pixel points around the pixel point respectively, and then determine the weighted average as the pixel value of the pixel point. Repeat the above operations until the resolution of the reconstructed image is the resolution determined according to the preset magnification.
[0142] In some embodiments, single-frame super-resolution reconstruction processing can be performed on the second sub-block image through a trained single-frame super-resolution reconstruction model.
[0143] In Figure 3In the illustrated embodiment, target pixel points are determined in the second sub-block image, the movement trajectories of the target pixel points are determined, and based on the movement trajectories of the target pixel points, the movement amplitudes between each image and the Kth frame image in the first K-1 frame images are determined. Then, based on the movement amplitudes between each image and the Kth frame image, a target movement amplitude is determined. When the target movement amplitude is less than or equal to a preset amplitude, it is determined that the super-resolution reconstruction method for the second sub-block image is a multi-frame super-resolution reconstruction method; when the target movement amplitude is greater than the preset amplitude, it is determined that the super-resolution reconstruction method for the second sub-block image is a single-frame super-resolution reconstruction method. Finally, according to the determined super-resolution reconstruction method, super-resolution reconstruction is performed on the second sub-block image to obtain reconstructed images of the second sub-block image at multiple resolutions. In the above method for performing super-resolution reconstruction on the second sub-block image, first, a suitable super-resolution reconstruction method is determined according to the target movement amplitude of the K frame images. When the target movement amplitude is less than or equal to the preset amplitude, multi-frame super-resolution reconstruction is used to perform super-resolution reconstruction on the second sub-block image; when the target movement amplitude is greater than the preset amplitude, single-frame super-resolution reconstruction is used to perform super-resolution reconstruction on the second sub-block image, so that after the super-resolution reconstruction of the second sub-block image, the obtained reconstructed images are clearer and have stronger information expression ability.
[0144] Based on the above embodiment, below, in combination with Figure 4 the illustrated embodiment, the process of performing face detection on the Kth frame image is further introduced.
[0145] Figure 4 The flowchart shows a process of performing face detection on the Kth frame image provided by an embodiment of the present application. Please refer to Figure 4 , which may include the following steps:
[0146] S41. Perform face detection processing on the Kth frame image to determine the first detection box in the Kth frame image.
[0147] The first detection box is the detection box obtained when performing face detection on the Kth frame image. For each first detection box, the first detection box corresponds to a label, which is used to indicate that the detection box is the first detection box obtained when performing face detection on the Kth frame image. In some embodiments, the first detection box may be represented by a rectangle or a square. Exemplarily, assuming K is 3, the detection boxes obtained when performing face detection on the 3rd frame image are detection box a and detection box b, then the first detection boxes in the Kth frame image are detection box a and detection box b, where the label corresponding to detection box a may be represented as 0, and the label corresponding to detection box b may be represented as 0.
[0148] In some embodiments, the Kth frame image may be input into a face detection model to obtain the first face detection box in the Kth frame image.
[0149] S42. For each second sub-block image, based on multiple reconstructed images of the second sub-block image, determine a second detection box corresponding to the second sub-block image in the K-th frame image.
[0150] The method for determining the second detection box corresponding to the second sub-block image in the K-th frame image based on multiple reconstructed images of the second sub-block image can be as follows: For each reconstructed image among the multiple reconstructed images, perform face detection processing on the reconstructed image to obtain a third detection box in the reconstructed image; based on the magnification of each reconstructed image relative to the second sub-block image, the position of each third detection box in the reconstructed image, and the position of the second sub-block image in the K-th frame image, determine the second detection box in the K-th frame image.
[0151] In some embodiments, the method for performing face detection processing on each reconstructed image among the multiple reconstructed images can be: input the reconstructed image into a face detection model to obtain a third detection box in the reconstructed image, where the third detection box is the detection box obtained after performing face detection processing on the reconstructed image. It should be noted that in some embodiments, to further improve the accuracy of face detection, face detection models reconstructed for different resolutions can be trained using reconstructed images at multiple resolutions, and then, during face detection, use the face detection models for reconstructed images of different resolutions to perform face detection on the reconstructed images of the corresponding resolutions. Exemplarily, assuming the resolutions are 200*200, 300*300, and 400*400, a face detection model 1 is trained using the reconstructed image with a resolution of 200*200, a face detection model 2 is trained using the reconstructed image with a resolution of 300*300, and a face detection model 3 is trained using the reconstructed image with a resolution of 400*400. During the face detection process, for the reconstructed image with a resolution of 200*200, the face detection model 1 can be used to perform face detection on it, for the reconstructed image with a resolution of 300*300, the face detection model 2 can be used to perform face detection on it, and for the reconstructed image with a resolution of 400*400, the face detection model 3 can be used to perform face detection on it.
[0152] In some embodiments, the third detection box can be represented by a rectangle or a square, etc. Taking the third detection box as a rectangle as an example, the position of the third detection box in the reconstructed image can be represented by the coordinates of the 4 corners of the third detection box, and the size of the third detection box can be represented by the length and width of the third detection box.
[0153] Exemplarily, assume that the third detection box is rectangular, and the four corners of the third detection box are X, Y, Z, and M respectively. Among them, the coordinates of corner X in the reconstructed image are (200, 10), the coordinates of corner Y in the reconstructed image are (300, 10), the coordinates of corner Z in the reconstructed image are (200, 70), and the coordinates of corner M in the reconstructed image are (300, 70). Then, the length of the third detection box is 100 and the width is 60. Therefore, the position of the third detection box in the reconstructed image can be expressed as [(200, 10), (300, 10), (200, 70), (300, 70)], and the size of the third detection box in the reconstructed image can be expressed as 100*60.
[0154] In some embodiments, for each third detection box, the third detection box corresponds to a label, which is used to represent the magnification of the reconstructed image where the third detection box is located relative to the second sub-block image. Exemplarily, the label corresponding to the third detection box can be expressed as n, where n represents the magnification of the reconstructed image where the third detection box is located relative to the second sub-block image.
[0155] In some embodiments, the second sub-block image can be represented by a rectangle or a square or other figures. Taking the second sub-block image as a rectangle as an example, the position of the second sub-block image in the Kth frame image can be expressed as the coordinates of the four corners of the second sub-block image in the Kth frame image. Exemplarily, assume that the four corners of the second sub-block image are Q, R, S, and T respectively. Among them, the coordinates of corner Q in the Kth frame image are (500, 200), the coordinates of corner R in the Kth frame image are (800, 200), the coordinates of corner S in the Kth frame image are (500, 300), and the coordinates of corner T in the Kth frame image are (800, 300). Therefore, the position of the second sub-block image in the Kth frame image can be expressed as [(500, 200), (800, 200), (500, 300), (800, 300)].
[0156] For any reconstructed image, the position and size of the detection box corresponding to the third detection box in the second sub-block image can be determined according to the magnification of the reconstructed image relative to the second sub-block image and the position of the third detection box in the reconstructed image. Then, according to the position and size of the detection box corresponding to the third detection box in the second sub-block image and the position of the second sub-block image in the Kth frame image, the second detection box is determined in the Kth frame image.
[0157] In some embodiments, the second detection box is the detection box corresponding to the third detection box in the Kth frame image. For each second detection box, the second detection box corresponds to a label, and it is consistent with the label of the third detection box corresponding to the second detection box and the label corresponding to the second detection box.
[0158] Exemplarily, assume that the position of the third detection box in the reconstructed image is [(200, 10), (300, 10), (200, 70), (300, 70)], the size of the third detection box in the reconstructed image can be expressed as 100*60, the label corresponding to the third detection box is 2, that is, the magnification of the reconstructed image relative to the second sub-block image is 2, then it can be determined that the position of the corresponding detection box of the third detection box in the second sub-block image is [(100, 5), (150, 5), (100, 35), (150, 35)], and the size of the third detection box in the second sub-block image is 50*60; assume that the position of the second sub-block image in the K-th frame image is [(500, 200), (800, 200), (500, 300), (800, 300)], then the position of the second detection box is determined in the K-th frame image as [(600, 205), (650, 205), (600, 235), (650, 235)], and the label corresponding to the second detection box is 2.
[0159] S43. According to the position of the first detection box in the K-th frame image, the confidence of the first detection box, the positions of the second detection boxes corresponding to the second sub-block images, and the confidence of the second detection boxes, perform screening processing on the first detection box and the second detection boxes corresponding to the second sub-block images to obtain the target detection box in the K-th frame image; wherein, the face detection result includes the position and size of the target detection box.
[0160] The confidence of the first detection box is used to represent the accuracy of including a face in the first detection box. In some embodiments, the confidence of the first detection box is a value between 0 and 1. The larger the value, the higher the accuracy of including a face in the first detection box; the smaller the value, the lower the accuracy of including a face in the first detection box. In some embodiments, the confidence of the first detection box can be a preset confidence.
[0161] Similarly, the confidence of the second detection box is used to represent the accuracy of including a face in the second detection box. In some embodiments, the confidence of the second detection box is a value between 0 and 1. The larger the value, the higher the accuracy of including a face in the second detection box; the smaller the value, the lower the accuracy of including a face in the second detection box. In some embodiments, according to the label of the second detection box, the magnification of the reconstructed image relative to the second sub-block image can be determined, and then the magnification of the reconstructed image relative to the second sub-block image is multiplied by the preset confidence to obtain the confidence of the second detection box.
[0162] Exemplarily, assume that the label of the second detection box is 2, then it is determined that the magnification of the reconstructed image relative to the second sub-block image is 2, and the preset confidence is 0.2, then the confidence of the second detection box is 2*0.2 = 4.
[0163] In some embodiments, the method for obtaining the target detection box in the K-th frame image by screening the first detection box and the second detection boxes corresponding to each second sub-block image may be as follows: For any one of the first detection boxes or the second detection boxes, determine at least one detection box that overlaps with this detection box. Respectively determine the overlapping areas between at least one detection box and this detection box. When there is a detection box whose overlapping area with this detection box is greater than the preset area, determine the confidence level of the detection box whose overlapping area with this detection box is greater than the preset area, as well as the confidence level of this detection box, and then determine the detection box with the highest confidence level as the target detection box. At this time, the position and size of the target detection box are the position and size of the detection box with the highest confidence level.
[0164] Exemplarily, the first detection boxes included in the K-th frame image are: first detection box A, first detection box B, and first detection box C, and the second detection boxes included in the K-th frame image are: second detection box D, second detection box E, and second detection box F. For the first detection box A, the detection boxes that overlap with the first detection box A are determined to be the first detection box C, the second detection box D, and the second detection box F. Among them, the overlapping area between the first detection box C and the first detection box A is 30, the overlapping area between the second detection box D and the first detection box A is 60, and the overlapping area between the second detection box F and the first detection box A is 100. The preset area is 50. Then, the detection boxes whose overlapping areas with the first detection box A are greater than the preset area are determined to be: the second detection box D and the second detection box F. The confidence level of the first detection box A is determined to be 0.4, the confidence level of the second detection box D is determined to be 0.7, and the confidence level of the second detection box F is determined to be 0.5. That is, the detection box with the highest confidence level is the second detection box D. Therefore, the second detection box D is determined as the target detection box.
[0165] In some embodiments, the non-maximum suppression (NMS) method may be used to screen the first detection box and the second detection boxes corresponding to each second sub-block image according to the position of the first detection box in the K-th frame image, the confidence level of the first detection box, the positions of the second detection boxes corresponding to each second sub-block image, and the confidence level of the second detection boxes, so as to obtain the target detection box in the K-th frame image.
[0166] In Figure 4In the illustrated embodiment, face detection is performed on the K-th frame image to obtain a first face detection box. Face detection is performed on multiple reconstructed images of each second sub-block image to obtain a third face detection box. Then, according to the position of the third face detection box in the reconstructed image, the magnification of the reconstructed image relative to the second sub-block image, and the position of the second sub-block image in the K-th frame image, a second detection box is determined in the K-th frame image. Finally, according to the position of the first detection box in the K-th frame image, the confidence level of the first detection box, the position of the second detection box in the K-th frame image, and the confidence level of the second detection box, the first detection box and the second detection box are screened to determine the target detection box, thereby obtaining the detection result of face detection on the K-th frame image. Through the above face detection method, face detection is respectively performed on the K-th frame image to obtain a first detection box, and face detection is performed on each reconstructed image to obtain a third detection box. According to the third detection box, a second detection box in the K-th frame image is determined, and the second detection box and the first detection box are screened to obtain the target detection box, thereby determining the detection result of face detection on the K-th frame image, and improving the accuracy of face detection on the K-th frame image.
[0167] Figure 5 The following is a schematic structural diagram of an image processing device provided by an embodiment of the present application. As Figure 5 shown, the image processing device 50 includes an acquisition module 51, a segmentation module 52, a reconstruction module 53, and a detection module 54, wherein,
[0168] The acquisition module 51 is configured to acquire K consecutive frame images, where K is a positive integer greater than 1;
[0169] The segmentation module 52 is configured to perform segmentation processing on the K frame images respectively to obtain a plurality of first sub-block images corresponding to each of the first K-1 frame images and a plurality of second sub-block images corresponding to the K-th frame image;
[0170] The reconstruction module 53 is configured to perform super-resolution reconstruction processing on each second sub-block image among the plurality of second sub-block images according to the plurality of first sub-block images corresponding to each of the first K-1 frame images to obtain a plurality of reconstructed images corresponding to the second sub-block image, and the resolutions of the plurality of reconstructed images are different;
[0171] The detection module 54 is configured to perform face detection on the K-th frame image according to the plurality of reconstructed images corresponding to each of the plurality of second sub-block images to obtain the face detection result of the K-th frame image.
[0172] In a possible implementation manner, the reconstruction module 53 is specifically configured to:
[0173] Determine the super-resolution reconstruction method of the second sub-block image according to the multiple first sub-block images corresponding to the first K-1 frame images respectively; the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method or a single-frame super-resolution reconstruction method;
[0174] Perform super-resolution reconstruction processing on the second sub-block image according to the super-resolution reconstruction method of the second sub-block image to obtain multiple reconstructed images corresponding to the second sub-block image.
[0175] In a possible implementation manner, the reconstruction module 53 is specifically configured to:
[0176] Determine target pixel points in the second sub-block image;
[0177] For each of the first K-1 frame images, perform matching processing on the multiple first sub-block images corresponding to the image and the second sub-block image, determine the target sub-block image in the multiple first sub-block images corresponding to the image, and the position of the target pixel point in the target sub-block image, where the target sub-block image is the first sub-block image to which the target pixel point belongs among the multiple first sub-block images corresponding to the image;
[0178] Determine the movement trajectory of the target pixel point according to the target sub-block images corresponding to the first K-1 frame images respectively and the positions of the target pixel point in the target sub-block images corresponding to the first K-1 frame images respectively;
[0179] Determine the super-resolution reconstruction method of the second sub-block image according to the movement trajectory of the target pixel point.
[0180] In a possible implementation manner, the reconstruction module 53 is specifically configured to:
[0181] Determine the movement amplitude between each image and the Kth frame image among the first K-1 frame images according to the movement trajectory of the target pixel point;
[0182] Determine the target movement amplitude according to the movement amplitudes between each image and the Kth frame image among the first K-1 frame images; the target movement amplitude is the average value of the movement amplitudes between each image and the Kth frame image among the first K-1 frame images, or the maximum value of the movement amplitudes between each image and the Kth frame image among the first K-1 frame images;
[0183] In the case where the target movement amplitude is less than or equal to the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is a multi-frame super-resolution reconstruction method;
[0184] In the case where the target movement amplitude is greater than the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is a single-frame super-resolution reconstruction method.
[0185] In a possible implementation manner, the reconstruction module 53 is specifically configured to:
[0186] Perform a fusion process on the second sub-block image and the target sub-block images corresponding to the previous K-1 frame images respectively to obtain a fused image corresponding to the second sub-block image;
[0187] Perform super-resolution reconstruction processing on the fused image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0188] In a possible implementation manner, the reconstruction module 53 is specifically configured to:
[0189] Perform super-resolution reconstruction on the second sub-block image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
[0190] In a possible implementation manner, the detection module 54 is specifically configured to:
[0191] Perform face detection processing on the Kth frame image to determine a first detection box in the Kth frame image;
[0192] For each second sub-block image, determine a second detection box corresponding to the second sub-block image in the Kth frame image according to multiple reconstructed images of the second sub-block image;
[0193] Perform a screening process on the first detection box and the second detection boxes corresponding to each second sub-block image according to the position of the first detection box in the Kth frame image, the confidence level of the first detection box, the positions of the second detection boxes corresponding to each second sub-block image, and the confidence levels of the second detection boxes to obtain a target detection box in the Kth frame image;
[0194] Wherein, the face detection result includes the position and size of the target detection box.
[0195] In a possible implementation manner, the detection module 54 is specifically configured to:
[0196] For each of the multiple reconstructed images, perform face detection processing on the reconstructed image to obtain a third detection box in the reconstructed image;
[0197] Determine a second detection box in the Kth frame image according to the magnification factor of each reconstructed image relative to the second sub-block image, the position of each third detection box in the reconstructed image, and the position of the second sub-block image in the Kth frame image.
[0198] The image processing device 50 provided in the embodiments of the present application can execute the technical solutions of the image processing method in the above method embodiments, and its implementation principles and beneficial effects are similar, which will not be elaborated here.
[0199] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 60 includes:
[0200] At least one processor 62; and
[0201] A memory 61 communicatively connected to the at least one processor 62; wherein
[0202] The memory 61 stores instructions executable by the at least one processor 62, and the instructions are executed by the at least one processor 62 to cause the at least one processor 62 to execute the image processing method involved in the above method embodiments.
[0203] Optionally, the above-mentioned processor may be a central processing unit (CPU), or may also be a GPU, other general-purpose processors, a digital signal processor (DSP), or an application specific integrated circuit (ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0204] The electronic device 60 provided in the embodiments of the present application can execute the image processing method involved in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.
[0205] A non-transitory computer-readable storage medium provided in the embodiments of the present application, wherein the computer-readable storage medium stores computer-executable instructions for causing a computer to execute the image processing method involved in the above method embodiments.
[0206] The embodiments of the present application provide a computer program product, including a computer program, which when executed by an electronic device, implements the image processing method involved in the above method embodiments.
[0207] The embodiments of the present application provide a chip, the chip includes at least one processor, and the processor is used to run program instructions to execute the image processing method in the above method embodiments.
[0208] The embodiments of the present application provide a chip module, on which a computer program is stored, and when the computer program is executed by the chip module, the image processing method in the above method embodiments is implemented.
[0209] All or part of the steps of the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable memory. When the program is executed, it performs the steps including the above method embodiments; and the foregoing memory (storage medium) includes: read-only memory (abbreviated as ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.
[0210] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processing machine, or other programmable terminal devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable terminal devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the function specified in one block or multiple blocks.
[0211] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a function specified in one block or multiple blocks.
[0212] These computer program instructions can also be loaded onto a computer or other programmable terminal devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the function specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a function specified in one block or multiple blocks.
[0213] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
[0214] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining K consecutive frames of images, where K is a positive integer greater than 1; Performing block processing on the K frames of images respectively to obtain a plurality of first block images corresponding to each of the first K-1 frames of images, and a plurality of second block images corresponding to the Kth frame of image; For each second block image among the plurality of second block images, performing super-resolution reconstruction processing on the second block image according to the plurality of first block images corresponding to the first K-1 frames of images respectively, to obtain a plurality of reconstructed images corresponding to the second block image, where the resolutions of the plurality of reconstructed images are different; Performing face detection on the Kth frame of image according to the plurality of reconstructed images corresponding to the plurality of second block images respectively, to obtain the face detection result of the Kth frame of image.
2. The method according to claim 1, wherein The performing super-resolution reconstruction processing on the second block image according to the plurality of first block images corresponding to the first K-1 frames of images respectively, to obtain a plurality of reconstructed images corresponding to the second block image, includes: Determining the super-resolution reconstruction method of the second block image according to the plurality of first block images corresponding to the first K-1 frames of images respectively; the super-resolution reconstruction method is a multi-frame super-resolution reconstruction method, or a single-frame super-resolution reconstruction method; Performing super-resolution reconstruction processing on the second block image according to the super-resolution reconstruction method of the second block image, to obtain a plurality of reconstructed images corresponding to the second block image.
3. The method according to claim 2, wherein The determining the super-resolution reconstruction method of the second block image according to the plurality of first block images corresponding to the first K-1 frames of images respectively, includes: Determining target pixel points in the second block image; For each frame of image among the first K-1 frames of images, performing matching processing on the plurality of first block images corresponding to the image and the second block image, to determine a target block image among the plurality of first block images corresponding to the image, and the position of the target pixel point in the target block image, where the target block image is the first block image to which the target pixel point belongs among the plurality of first block images corresponding to the image; Determining the movement trajectory of the target pixel point according to the target block images corresponding to the first K-1 frames of images respectively, and the position of the target pixel point in the target block images corresponding to the first K-1 frames of images respectively; Determining the super-resolution reconstruction method of the second block image according to the movement trajectory of the target pixel point.
4. The method according to claim 3, wherein The determining the super-resolution reconstruction method of the second block image according to the movement trajectory of the target pixel point, includes: Determining the movement amplitude between each image and the Kth frame of image in the first K-1 frames of images according to the movement trajectory of the target pixel point; Determining a target movement amplitude according to the movement amplitude between each image and the Kth frame of image in the first K-1 frames of images; the target movement amplitude is the average value of the movement amplitudes between each image and the Kth frame of image in the first K-1 frames of images, or the maximum value of the movement amplitudes between each image and the Kth frame of image in the first K-1 frames of images; When the target movement amplitude is less than or equal to the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the multi-frame super-resolution reconstruction method; When the target movement amplitude is greater than the preset amplitude, determine that the super-resolution reconstruction method of the second sub-block image is the single-frame super-resolution reconstruction method.
5. The method according to claim 3 or 4, characterized in that, The super-resolution reconstruction method is the multi-frame super-resolution reconstruction method. According to the super-resolution reconstruction method of the second sub-block image, performing super-resolution reconstruction processing on the second sub-block image to obtain multiple reconstructed images corresponding to the second sub-block image includes: Performing fusion processing on the second sub-block image and the target sub-block images corresponding to the previous K-1 frame images respectively to obtain a fused image corresponding to the second sub-block image; Performing super-resolution reconstruction processing on the fused image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
6. The method according to any one of claims 2-4, characterized in that, The super-resolution reconstruction method is the single-frame super-resolution reconstruction method. According to the super-resolution reconstruction method of the second sub-block image, performing super-resolution reconstruction processing on the second sub-block image to obtain multiple reconstructed images corresponding to the second sub-block image includes: Performing super-resolution reconstruction on the second sub-block image according to multiple preset magnification factors to obtain reconstructed images of the second sub-block image at multiple resolutions.
7. The method according to any one of claims 1-4, characterized in that, According to the multiple reconstructed images corresponding to the multiple second sub-block images respectively, performing face detection on the Kth frame image to obtain the face detection result of the Kth frame image includes: Performing face detection processing on the Kth frame image to determine the first detection box in the Kth frame image; For each of the second sub-block images, determining a second detection box corresponding to the second sub-block image in the Kth frame image according to the multiple reconstructed images of the second sub-block image; According to the position of the first detection box in the Kth frame image, the confidence of the first detection box, the positions of the second detection boxes corresponding to the second sub-block images, and the confidence of the second detection boxes, performing screening processing on the first detection box and the second detection boxes corresponding to the second sub-block images to obtain the target detection box in the Kth frame image; Wherein, the face detection result includes the position and size of the target detection box.
8. The method according to claim 7, characterized in that Determining the second detection box corresponding to the second sub-block image in the Kth frame image according to the multiple reconstructed images of the second sub-block image includes: For each of the multiple reconstructed images, performing face detection processing on the reconstructed image to obtain a third detection box in the reconstructed image; According to the magnification factor of each reconstructed image relative to the second sub-block image, the position of each third detection box in the reconstructed image, and the position of the second sub-block image in the Kth frame image, determining the second detection box in the Kth frame image.
9. An image processing apparatus, characterized in that, Including: An acquisition module, configured to acquire K consecutive frame images, where K is a positive integer greater than 1; A chunking module, configured to perform chunking processing on the K frame images respectively, to obtain a plurality of first chunked images corresponding to each of the first K-1 frame images, and a plurality of second chunked images corresponding to the Kth frame image; A reconstruction module, configured to perform super-resolution reconstruction processing on each of the second chunked images among the plurality of second chunked images, according to the plurality of first chunked images corresponding to each of the first K-1 frame images, to obtain a plurality of reconstructed images corresponding to the second chunked image, where the resolutions of the plurality of reconstructed images are different; A detection module, configured to perform face detection on the Kth frame image according to the plurality of reconstructed images corresponding to each of the plurality of second chunked images, to obtain a face detection result of the Kth frame image.
10. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor executes the image processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Preview image repair method and device, electronic equipment and computer program product
CN122530034A