Portrait file processing method and device, storage medium and electronic equipment

By classifying and image repairing the capture data queue, the problem of poor portrait aggregation effect caused by the loss of capture data is solved, the utilization rate of capture data and portrait aggregation effect are improved, and the effective utilization and clustering of high-quality images are achieved.

CN120375022APending Publication Date: 2025-07-25ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310995.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, due to the poor aggregation effect of portraits caused by the loss of capture data, the existing methods lose capture information when discarding low-quality images, affecting the portrait clustering effect.

Method used

By classifying the capture data queue, the images to be clustered and to be repaired are selected, and the image repair model is used to determine the reference image based on the capture time and point to perform image repair, and the repaired images are reflowed and clustered. The high-quality images after blurring are used as input and output examples of the image repair model, without the need for fine-tuning of the model.

Benefits of technology

It improves the utilization rate of capture data and portrait aggregation effect, improves the adaptability of image repair models, and ensures the utilization and clustering effect of high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375022A_ABST
    Figure CN120375022A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a portrait file processing method and device, a storage medium and electronic equipment, and the method comprises the steps: carrying out the classification processing of snapshot data in a snapshot data queue, and obtaining a group of to-be-clustered images and a group of to-be-restored images; determining a reference image corresponding to each to-be-restored image in the group of to-be-restored images from the group of to-be-clustered images based on the snapshot time and the snapshot point location; inputting each to-be-restored image and the reference image corresponding to each to-be-restored image into an image restoration model to obtain a restored image corresponding to each to-be-restored image; and performing portrait clustering operation on the group of to-be-clustered images and the repaired images meeting the preset clustering condition to obtain a group of portrait archives. Through the method and the device, the technical problem of poor portrait aggregation effect caused by loss of the snapshot data in a processing method of portrait file processing in the related technology is solved, and the effects of improving the utilization rate of the snapshot data and improving the portrait aggregation effect are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image clustering, and more specifically, to a method and device for processing portrait archives, a storage medium, and an electronic device. Background Art

[0002] Due to the complexity of the real environment and the randomness of point capture, the collection of personnel images often results in some low-quality images (i.e., blurred images, abnormal exposure, noise, etc.). For example, if there are people moving at high speed in front of the capture point, the captured images will often appear blurred; at certain angles at noon, the captured images of people will often be overexposed, resulting in reduced image recognizability; when the light is insufficient, the captured images will produce serious image noise.

[0003] In the related art, several evaluation indicators are usually selected and a fixed threshold is assigned to each evaluation indicator. When the captured image does not meet the threshold of a certain indicator, it is directly discarded and does not participate in the clustering task. However, although the above solution solves the computing resource overhead to a certain extent, it also loses the captured information of the person, affecting the improvement of the effect.

[0004] It can be seen that the portrait file processing method in the related art has the problem of poor portrait aggregation effect due to the loss of captured data. Summary of the invention

[0005] The embodiments of the present application provide a method and device for processing a portrait file, a storage medium, and an electronic device, so as to at least solve the technical problem of poor portrait aggregation effect caused by the loss of captured data in the processing method of portrait file processing in the related art.

[0006] According to one aspect of an embodiment of the present application, a method for processing a portrait file is provided, comprising:

[0007] Classifying and processing the captured data in the captured data queue to obtain a group of images to be clustered and a group of images to be restored, wherein the captured data queue is generated based on a group of captured data of a captured device, the group of images to be clustered is captured data that meets a preset clustering condition, and the group of images to be restored is captured data that does not meet the preset clustering condition, and the preset clustering condition is a condition that is satisfied by the images that allow portrait clustering;

[0008] Based on the capture time and the capture point, determining a reference image corresponding to each image to be restored in the group of images to be restored from the group of images to be clustered;

[0009] Input each of the images to be repaired and the reference image corresponding to each of the images to be repaired into an image inpainting model to obtain the inpainted image corresponding to each of the images to be repaired, where the image inpainting model is a model for inpainting the input image to be repaired by using the blurred reference image and the reference image as input and output image examples;

[0010] Perform a portrait clustering operation on the set of images to be clustered and the inpainted images that meet the preset clustering conditions to obtain a set of portrait profiles.

[0011] According to another aspect of the embodiments of the present application, there is also provided a processing device for portrait profiles, including:

[0012] A classification unit for classifying the capture data in the capture data queue to obtain a set of images to be clustered and a set of images to be repaired, where the capture data queue is generated based on the capture data of a set of capture devices, the set of images to be clustered is the capture data that meets the preset clustering conditions, the set of images to be repaired is the capture data that does not meet the preset clustering conditions, and the preset clustering conditions are the conditions satisfied by the images that allow portrait clustering;

[0013] A determination unit for determining, based on the capture time and the capture location, the reference image corresponding to each of the images to be repaired in the set of images to be repaired from the set of images to be clustered;

[0014] An input unit for inputting each of the images to be repaired and the reference image corresponding to each of the images to be repaired into an image inpainting model to obtain the inpainted image corresponding to each of the images to be repaired, where the image inpainting model is a model for inpainting the input image to be repaired by using the blurred reference image and the reference image as input and output image examples;

[0015] A clustering unit for performing a portrait clustering operation on the set of images to be clustered and the inpainted images that meet the preset clustering conditions to obtain a set of portrait profiles.

[0016] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0017] According to another aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.

[0018] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the steps in any one of the above method embodiments through the computer program.

[0019] Through the present application, a waste film reflux mechanism is adopted to classify the captured data in the captured data queue, and a set of images to be clustered and a set of images to be repaired are obtained. Among them, the captured data queue is generated based on the captured data of a set of capture devices. A set of images to be clustered are captured data that meet the preset clustering conditions, and a set of images to be repaired are captured data that do not meet the preset clustering conditions. The preset clustering conditions are the conditions satisfied by images that allow portrait clustering. Based on the capture time and the capture location, a reference image corresponding to each image to be repaired in the set of images to be repaired is determined from the set of images to be clustered. Each image to be repaired and the reference image corresponding to each image to be repaired are input into an image repair model, and a repaired image corresponding to each image to be repaired is obtained. Among them, the image repair model is a model that uses the blurred reference image and the reference image as input and output image examples to perform image repair on the input image to be repaired. A portrait clustering operation is performed on the set of images to be clustered and the repaired images that meet the preset clustering conditions to obtain a set of portrait files. Since the images to be repaired (belonging to waste film data, that is, images that do not meet the preset clustering conditions) are repaired and the repaired images that meet the preset clustering conditions are refluxed for clustering, the purpose of improving the utilization rate of captured data can be achieved, and the technical effect of improving the portrait aggregation effect can be achieved. Furthermore, the technical problem of poor portrait aggregation effect caused by the loss of captured data in the processing method of portrait files in the related art can be solved. In addition, since high-quality images (high-quality data, that is, images that meet the preset clustering conditions) after blurring processing and high-quality images are used as input and output image examples of the image repair model, the image repair model can be adapted to the current image to be repaired without model fine-tuning or any other model modification, which can improve the adaptability of the model and the effect of image repair. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of an application scenario of a method for processing portrait files according to an embodiment of the present application;

[0021] Figure 2 It is a schematic flowchart of an optional method for processing portrait files according to an embodiment of the present application;

[0022] Figure 3 It is a schematic flowchart of another optional method for processing portrait files according to an embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of an optional method for processing portrait files according to an embodiment of the present application;

[0024] Figure 5 It is a schematic diagram of another optional method for processing portrait files according to an embodiment of the present application;

[0025] Figure 6 It is a schematic diagram of yet another optional method for processing portrait files according to an embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of yet another optional method for processing portrait files according to an embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of yet another optional method for processing portrait files according to an embodiment of the present application;

[0028] Figure 9 It is a schematic diagram of yet another optional method for processing portrait files according to an embodiment of the present application;

[0029] Figure 10 It is a schematic diagram of yet another optional method for processing portrait files according to an embodiment of the present application;

[0030] Figure 11 It is a schematic flowchart of yet another optional method for processing portrait files according to an embodiment of the present application;

[0031] Figure 12 It is a structural block diagram of an optional device for processing portrait files according to an embodiment of the present application;

[0032] Figure 13 It is a structural block diagram of a computer system of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0033] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0034] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] According to one aspect of the embodiments of this application, a method for processing portrait files is provided. Optionally, in this embodiment, the above method for processing portrait files may but is not limited to be applied to a hardware environment including a capture device 102 and a server 104 as shown in Figure 1 Figure. The server 104 can be connected to the capture device 102 through a network, and can be used to provide services (such as application services, etc.) for the capture device 102 or the client installed on the capture device 102. A database can be set up on the server 104 or independently of the server 104 to provide data storage services for the server 104.

[0036] The above network may include but is not limited to at least one of the following: wired network, wireless network. The above wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The capture device 102 may but is not limited to be a monocular camera, a binocular camera, a camera, etc. The server 104 may but is not limited to be a cloud server, a server cluster or other server types.

[0037] The method for processing portrait files in the embodiments of this application can be executed by the server 104, or can be jointly executed by the server 104 and the capture device 102. Among them, when the capture device 102 executes the method for processing portrait files in the embodiments of this application, it can also be executed by the client installed on it.

[0038] Taking the processing method of the portrait file in this embodiment executed by the server 104 as an example, Figure 2 It is a schematic flowchart of an optional processing method of a portrait file according to an embodiment of the present application. As Figure 2 shown, the process of this method may include the following steps:

[0039] Step S202: Classify the captured data in the captured data queue to obtain a set of images to be clustered and a set of images to be repaired. Among them, the captured data queue is generated based on the captured data of a set of capture devices. A set of images to be clustered is captured data that meets the preset clustering conditions, and a set of images to be repaired is captured data that does not meet the preset clustering conditions. The preset clustering conditions are the conditions that the images allowing portrait clustering should meet;

[0040] Step S204: Based on the capture time and capture location, determine the reference image corresponding to each image to be repaired in a set of images to be repaired from a set of images to be clustered;

[0041] Step S206: Input each image to be repaired and the reference image corresponding to each image to be repaired into the image repair model to obtain the repaired image corresponding to each image to be repaired. Among them, the image repair model is a model used to perform image repair on the input image to be repaired with the blurred reference image and the reference image as the input and output image examples;

[0042] Step S208: Perform a portrait clustering operation on a set of images to be clustered and the repaired images that meet the preset clustering conditions to obtain a set of portrait files.

[0043] The processing method of the portrait file in this embodiment can be applied to the field of image clustering and applied to the scenario of portrait clustering. Here, portrait clustering refers to the process of analyzing and grouping the captured images at the points. Considering that the capture is usually of human-shaped objects, the captured images can also be called personnel images. Due to the complexity of the real environment and the randomness of point capture, the image acquisition of personnel often shows some low-quality images (that is, images with image blurring, abnormal exposure, noise, etc.). For example, when there are people moving at high speed in front of the capture point, the captured images are often blurred; at noon at certain angles, the captured images are often overexposed, resulting in a decrease in the image recognition rate; when the light is insufficient, the captured images will generate serious image noise. Therefore, in order to ensure the quality of the pictures entering the subsequent processing process, the captured images obtained can be initially screened before portrait clustering.

[0044] It should be noted that in this embodiment, unless otherwise specified, capture data, capture images, and capture pictures are the same concept described from different perspectives; waste film data, low-quality images, and low-quality pictures are the same concept described from different perspectives; high-quality data, high-quality images, high-quality pictures, and images to be clustered are the same concept described from different perspectives; images to be repaired and pictures to be repaired are the same concept described from different perspectives. Among them, the image to be repaired can be all or part of the waste film data.

[0045] In the related art, the preliminary screening of capture images is usually performed based on a series of evaluation indicators (such as image sharpness, exposure, and noise level, etc.). These evaluation indicators are given fixed thresholds. When a capture image does not meet the threshold of a certain evaluation indicator, it is directly discarded and does not participate in the subsequent clustering task. Although the above image screening method can reduce the overhead of computing resources to a certain extent, it will also lose the capture information of personnel and affect the improvement of the portrait clustering effect.

[0046] To at least partially solve the above technical problems, a waste film return mechanism can be adopted. A low-quality image repair module is introduced into the portrait clustering system to repair low-quality images (i.e., waste film data), and the repaired images are returned for clustering to improve the utilization rate of capture data: repairing the images to be repaired in the low-quality images and returning the repaired images that meet the preset clustering conditions for clustering can achieve the purpose of improving the utilization rate of capture data and improving the portrait aggregation effect; in addition, using the blurred high-quality images and high-quality images as the input and output image examples of the image repair model can make the image repair model adapt to the current image to be repaired without model fine-tuning or any other model modification, which can improve the adaptability of the model and the effect of image repair. Here, the waste film data refers to the capture images that do not participate in the clustering task due to the portrait deflection angle, pitch angle, blur, etc. not meeting the clustering rules of the current portrait clustering system.

[0047] In this embodiment, a set of capture devices can be set. The number of capture devices in a set of capture devices can be one or more. To improve the comprehensiveness of portrait file establishment, the number of capture devices in a set of capture devices is usually multiple, and the capture points where different capture devices are set can be different. For the server, it can continuously obtain capture data from a set of capture devices and put the capture data directly or after certain processing into the capture data queue. Here, the capture data queue is generated based on the capture data of a set of capture devices, and the capture data in the capture data queue can be the capture data directly obtained from the capture devices or the capture data after certain processing (such as image preprocessing, image cropping, etc.).

[0048] Considering the differences among the characteristics of different human body parts, in order to improve the accuracy of portrait clustering, one or more human body parts can be selected as the targets of portrait clustering. Correspondingly, the number of captured data queues can be one or more, and one captured data queue can correspond to one human body part. In addition, it is also possible to directly put the captured data obtained from each capturing device into a single captured data queue without distinguishing human body parts, and this embodiment does not limit this.

[0049] To facilitate portrait clustering, conditions that the images allowed for portrait clustering need to meet can be set, that is, preset clustering conditions. For the captured data in the captured data queue, it can be determined whether each of them meets the preset clustering conditions, and the captured data in the captured data queue can be classified based on the preset clustering conditions. The results of the classification process can include: a set of images to be clustered and a set of images to be repaired. A set of images to be clustered are the captured data that meet the preset clustering conditions, while a set of images to be repaired are the captured data that do not meet the above preset clustering conditions. Here, a set of images to be repaired can be all the captured data that do not meet the above preset clustering conditions, or can be selected from the captured data that do not meet the above preset clustering conditions with reference to the conditions allowing image repair. Optionally, the preset clustering conditions can include, but are not limited to, conditions set based on at least one of the following parameters: yaw angle, pitch angle, image quality parameters (such as image sharpness, exposure, color saturation, etc.).

[0050] The classification process of the captured data can be performed based on the image features of the different captured data parsed and / or the attribute values of a set of specified attributes. Here, the image features are the vectorized representations of the images, that is, the feature vectors of the images, and each specified attribute can correspond to a characteristic of the humanoid object in the image, such as face width, whether wearing a mask, etc. The parsing process of the captured data can be executed by a parsing module, and the parsing module can be used to perform feature extraction, attribute value parsing, etc. on the captured data, and convert the captured images into the data forms required by the portrait clustering system.

[0051] For each image to be repaired, multiple methods can be used for image repair. For example, an image repair model based on a generative adversarial network can be used for image repair. However, the image repair model based on a generative adversarial network relies on a large batch of datasets to train the repair model, and there are certain limitations in practical applications. Another example is that an image repair method based on a multimodal large model can be adopted. However, the above image repair method first locates, extracts features, and performs feature space transformation on the image to be repaired, and then inputs it into the multimodal large model, which requires continuous parameter adjustment, and the usability and systematicness are relatively low.

[0052] It can be seen that in the image restoration model in the related art, a large amount of data pre-training is required and the model parameters are fine-tuned to be applicable to the required application scenarios. However, new application scenarios often do not have a large amount of data to train large and complex network models, which limits the wide application of large network models. In this embodiment, the blurred high-quality image and the high-quality image are used as input and output image examples of the image restoration model, and the image restoration model can be adapted to the current image to be restored without model fine-tuning or any other model modification, which can improve the adaptability of the model and the effect of image restoration.

[0053] After obtaining a set of images to be clustered and a set of images to be restored, based on the capture time and capture location, a reference image corresponding to each image to be restored in the set of images to be clustered can be determined. Here, the reference image corresponding to each image to be restored is the image to be clustered used for image restoration of each image to be restored, and the image to be clustered is the aforementioned high-quality image or high-quality data. Screening the reference images based on the capture time and capture location can ensure the correlation between the selected reference image and the corresponding image to be restored, thereby improving the effect of image restoration.

[0054] The image restoration model used for image restoration in this embodiment is a model that uses the blurred reference image and the reference image as input and output image examples (image Prompt) to perform image restoration on the input image to be restored. Here, the image Prompt enables the model to execute multiple specified tasks by providing new input and output image examples to the pre-trained model without model fine-tuning or any other model modification. For each image to be restored, each image to be restored and the reference image corresponding to each image to be restored can be input into the image restoration model. The image restoration model can blur the reference image corresponding to each image to be restored, and use the blurred reference image corresponding to each image to be restored and the reference image corresponding to each image to be restored as input and output image examples. Referring to the conversion relationship between the blurred reference image corresponding to each image to be restored and each image to be restored, image restoration is performed on each image to be restored to obtain the restored image corresponding to each image to be restored.

[0055] The restored image corresponding to each image to be restored may have a restored image that meets the preset clustering conditions, or may have a restored image that does not meet the preset clustering conditions. And portrait clustering is performed on the images that meet the preset clustering conditions. Therefore, portrait clustering operations can be performed on a set of images to be clustered and the restored images that meet the preset clustering conditions to obtain a set of portrait files. Here, the portrait file is a set generated by portrait clustering, and all the captured images of one natural person can be put into one set.

[0056] Exemplarily, in the field of portrait clustering, the personnel trajectory can be appended based on the visual large model image restoration. By using the large model for image restoration, an image restoration function module relying on "image Prompt" is constructed to realize the restoration and recycling clustering of waste film data, so as to improve the utilization rate of captured data. Among them, the appended personnel trajectory can be applied to the process of portrait clustering, and the large model used by the image restoration function module is a low-quality image restoration model (i.e., the image restoration model). Combined with Figure 3 , the method flow of portrait clustering includes the following steps S302 to S308.

[0057] Step S302, parse the captured image and perform waste film verification. Parsing the captured image can be executed by the aforementioned parsing module. Here, referring to Figure 4 , the captured image can be parsed, and the parsed feature values and attribute information (attribute values) are output. That is, the parsing results of each captured image can include their respective corresponding feature values and various attribute values. Based on the parsing results of the captured image, waste film verification is performed, that is, the quality of the captured image is evaluated, and waste film data and high-quality data are classified and output. The waste film verification can be executed by the waste film verification module.

[0058] Step S304, the waste film learning model outputs a high-quality image (i.e., the reference image) corresponding to the low-quality image (i.e., the waste film data to be restored). Here, the waste film learning model can select the corresponding high-quality image based on the image features of the low-quality image, and the image features of the low-quality image can be the parsing results of the low-quality image. Referring to Figure 5 , the feature information and attribute information of the parsed captured image are obtained, and the feature information and attribute information of the captured image are evaluated for quality. Based on the evaluation results, waste film data (i.e., low-quality images) and high-quality data (i.e., high-quality images) are output. The output waste film data and high-quality data can be input into the waste film learning model, and the waste film learning model selects the high-quality data corresponding to the waste film data.

[0059] S306, the low-quality image restoration model (i.e., the image restoration model, which can be a large model) performs image restoration on the low-quality image. The low-quality image and the high-quality image corresponding to the low-quality image can be input into the low-quality image restoration model, and the low-quality image restoration model outputs the low-quality image after image restoration.

[0060] S308, the high-quality image participates in portrait clustering. The low-quality image may be a high-quality image or a low-quality image after image restoration. Whether it is the high-quality image classified and output or the high-quality image restored, both can participate in portrait clustering to obtain the aggregated portrait file.

[0061] Among them, the low-quality images that may be repaired in the waste film data can be put into the waste film message queue. For the low-quality image repair model, as Figure 6 shown, the waste film data in the waste film message queue and the high-quality data that pass the waste film verification are output to the large low-quality image repair model. The large low-quality image repair model can repair the input waste film data based on the input high-quality data and output the repaired data.

[0062] The low-quality image repair model can be composed of NLP (Natural Language Processing) plus an image repair network model, and it can be a kind of LVM (Large Vision Models). In the existing image repair models, a large amount of data pre-training is often required, and the model parameters are fine-tuned to be applicable to the scenarios that need to be applied. However, new application scenarios often do not have a large amount of data to train large and complex network models, which limits the wide application of large network models. The low-quality image repair model can use "Prompt" to achieve the operation of a similar network model on images. By using the high-quality data that is consistent with the source of the waste film data and passes the waste film verification as an "image Prompt" to the network model, and using the form of NLP plus an image repair model, the repair of the aggregated waste film data can be completed, and the pre-trained model can be applied to a new scenario without specific task fine-tuning and any model modification. Among them, the input data of the low-quality image repair model is: high-quality data, output examples, waste film data, and the output data is the repaired data.

[0063] First, a network structure can be constructed, and then the constructed network structure is input into the large image repair model (Impainting model), and the repaired image output by the large image repair model can be obtained. As Figure 7 shown, the constructed network structure ( Figure 7 the input part in) includes the blurred reference image in the left column and the reference image (output example) in the right column. The image to be repaired is the image at the bottom of the left column, and the image to be generated is the image at the bottom of the right column. The reference image can be the clustering capture image to be verified (high-quality image) that has passed the waste film verification link. To ensure the image repair effect, there is a certain relationship between the high-quality image as the reference image and the image to be repaired. The blurred reference image can be obtained by performing a Gaussian blur operation on the reference image, and the blurred reference image is used as a reference to supplement information for the image to be repaired, and the repair data of the image to be repaired is output.

[0064] Here, taking advantage of the superiority of the deep network model in image restoration, the data with poor quality of the front-end captured images can be restored again, which can reduce the loss of personnel trajectories, help improve the image recall rate, the point position recall rate, and the effect of portrait clustering, and at the same time improve the utilization rate of the captured data.

[0065] Through the embodiments provided in this application, the captured data in the captured data queue is classified to obtain a set of images to be clustered and a set of images to be restored. Among them, the captured data queue is generated based on the captured data of a set of capturing devices. A set of images to be clustered are the captured data that meet the preset clustering conditions, and a set of images to be restored are the captured data that do not meet the preset clustering conditions. The preset clustering conditions are the conditions that the images allowing portrait clustering satisfy; based on the capture time and the capture point position, a reference image corresponding to each image to be restored in the set of images to be restored is determined from the set of images to be clustered; each image to be restored and the reference image corresponding to each image to be restored are input into the image restoration model to obtain the restored image corresponding to each image to be restored. Among them, the image restoration model is a model that uses the blurred reference image and the reference image as the input and output image examples to perform image restoration on the input image to be restored; a portrait clustering operation is performed on the set of images to be clustered and the restored images that meet the preset clustering conditions to obtain a set of portrait files, which solves the problem of poor portrait aggregation effect caused by the loss of captured data in the portrait file processing method in the related technology, improves the utilization rate of the captured data, and enhances the portrait aggregation effect.

[0066] In an exemplary embodiment, the waste film data may be the captured images that do not participate in the clustering task because the portrait deflection angle, pitch angle, blur, etc. do not conform to the clustering rules of the current portrait clustering system. When performing classification processing, each captured data can be evaluated based on a set of evaluation parameters, and image classification is performed based on the evaluation results. Here, the set of evaluation parameters includes an image quality parameter and at least one of the following: deflection angle, pitch angle. Correspondingly, classifying the captured data in the captured data queue to obtain a set of images to be clustered and a set of images to be restored includes: determining the parameter values corresponding to each captured data in the captured data queue and each evaluation parameter in the set of evaluation parameters; in the captured data queue, the captured data whose parameter values corresponding to each evaluation parameter are all within the parameter value range of each evaluation parameter are determined as the images to be clustered, and a set of images to be clustered is obtained; in the captured data queue, the captured data whose parameter value corresponding to the image quality parameter is not within the parameter value range of the image quality parameter and whose parameter values corresponding to the other evaluation parameters except the image quality parameter in the set of evaluation parameters are all within the parameter value range of the other evaluation parameters are determined as the images to be restored, and a set of images to be restored is obtained.

[0067] Determining the parameter values corresponding to each captured data and each evaluation parameter can be performed by the aforementioned parsing module, or after the parsing module parses out the parameter values corresponding to each captured data and each evaluation parameter, it can be passed to the module that performs the classification process. For each captured data, if the parameter values corresponding to each captured data and each evaluation parameter are all within the parameter value range of each evaluation parameter, it can be determined as an image to be clustered (high-quality data, that is, high-quality portrait data). If the parameter values corresponding to it and at least one evaluation parameter are not within the corresponding parameter value range, it can be determined as waste film data. Through the above processing, a set of images to be clustered and a set of waste film data can be obtained.

[0068] All waste film data can be used as images to be repaired. Optionally, considering that problems such as the deflection angle and the pitch angle usually cannot be overcome by image repair, in order to reduce the workload of image repair, only the waste film data that does not meet the requirements due to image quality parameters can be repaired. Correspondingly, in the captured data queue, the captured data whose parameter values corresponding to the image quality parameters are not within the parameter value range of the image quality parameters and whose parameter values corresponding to other evaluation parameters except the image quality parameters in a set of evaluation parameters are all within the parameter value range of other evaluation parameters can be determined as images to be repaired, and a set of images to be repaired is obtained.

[0069] For example, data classification can be performed according to whether the index values (that is, the parameter values of the evaluation parameters) of evaluation indexes such as the deflection angle, the pitch angle, and qescore (that is, the picture quality evaluation value) are within the specified range, and finally two types of data are obtained: waste film data and high-quality data. Waste film data is generally caused by various reasons and corresponds to multiple waste film type identifiers. Only the data that enters the waste film table due to qescore not meeting the regulations can be repaired. When qescore does not meet the value range, the intuitive feeling of the picture quality is that the picture is blurred and the people in the picture are not clear, and such waste films can be repaired.

[0070] Here, the waste film table is a data table that records waste film data. The waste film type refers to the reason for entering the waste film table. For example, the deflection angle is greater than the specified range, the picture quality value is lower than the specified range, etc. Qescore refers to the total score of picture quality evaluation. The higher the score, the better the image quality.

[0071] Through this embodiment, image screening based on image quality parameters and at least one of the deflection angle and the pitch angle can improve the efficiency of image screening; only using the images that do not meet the requirements of image quality as images to be repaired can reduce the amount of data to be processed for image repair and improve the efficiency of image repair.

[0072] In an exemplary embodiment, based on the large image inpainting model, a mechanism of using high-quality captures at the same position within a neighboring unit time of the waste film data as the "image Prompt" is implemented to achieve the implementation of the adaptability of the large image inpainting model in a specific application scenario. Correspondingly, based on the capture time and capture position, a reference image corresponding to each to-be-restored image in a set of to-be-clustered images is determined, including: taking each to-be-restored image as the current to-be-restored image and performing the following processing operations to obtain a reference image corresponding to each to-be-restored image: in the case where there are to-be-clustered images that meet the preset image conditions in a set of to-be-clustered images, at least one to-be-clustered image that meets the preset image conditions is determined as the reference image corresponding to the current to-be-restored image.

[0073] For the obtained set of to-be-restored images, each to-be-restored image in the set of to-be-clustered images can be taken as the current to-be-restored image for reference image selection to obtain a reference image corresponding to each to-be-restored image. Based on the current to-be-restored image, it can be determined whether there are to-be-clustered images that meet the preset image conditions in the set of to-be-clustered images. If so, at least one to-be-clustered image that meets the preset image conditions is determined as the reference image corresponding to the current to-be-restored image.

[0074] Here, the preset image conditions are used as the standardized conditions for screening reference images to ensure the relevance of the reference images and the to-be-restored images in terms of time, position, and angle, thereby improving the accuracy and effect of restoration. The preset image conditions can be that the time difference between the capture time and the capture time of the current to-be-restored image is less than or equal to the preset time difference threshold, the capture position is the same as the capture position of the current to-be-restored image, and the angle difference between the angle value of the specified angle and the angle value of the specified angle of the current to-be-restored image is within the preset angle difference range. The specified angle includes at least one of the following: yaw angle, pitch angle. The judgment based on the capture time, capture position, and specified angle can be executed serially or in parallel, and the order of judgment can be adjusted as needed. This is not limited in this embodiment.

[0075] It should be noted that the preset time difference threshold can be used to measure the proximity in time between the current to-be-restored image and the to-be-clustered images, and its value can be 2 seconds, 5 seconds, 10 seconds, or other values, which can be flexibly adjusted according to factors such as hardware resources and data volume. For example, if the preset time difference threshold is 10 seconds, the system will look for high-quality images captured within 10 seconds in the to-be-clustered images as references. As Figure 8 shown, the capture Figure 1 is the to-be-restored image, and the corresponding reference image is selected from the captured images within 10 seconds before and after the to-be-restored image. The capture Figure 2 , the capture Figure 3 and the capture Figure 4 are possible reference images.

[0076] For the capture point, the selected reference image can be the same as the capture point of the current image to be repaired, so as to ensure that the capture positions of the reference image and the current image to be repaired are the same (maintaining the consistency of the human background) and avoid the background mismatch problem during the repair process.

[0077] For the angle value of the specified angle, the angular difference between the selected reference image and the angle value of the specified angle of the current image to be repaired is within the preset angular difference range. The specified angle includes at least one of the following: yaw angle, pitch angle. The preset angular difference range is the set allowable angular deviation, which is used to evaluate the similarity of the current image to be repaired and the candidate reference image in terms of the human pose angle. The preset angular difference range can be a fixed range. For example, [0~0.5], [0.1~0.5], [1~5], [3~6], etc. Taking [0~0.5] as an example, the images to be clustered with an angular difference (for example, the angular difference between yaw angles, the angular difference between pitch angles) not exceeding 0.5 degrees from the angle of the image to be repaired can be used as reference images.

[0078] For example, there is a certain relationship between the reference image and the image to be repaired. Select the images to be clustered with consistent point position information, an angular range between [0.1~0.5], and passing the waste film verification process from the captured images captured 10 seconds before and after the image to be repaired as the reference image of the image to be repaired.

[0079] Exemplarily, the capture time of the image to be repaired is 14:02:15, the capture point is entrance A, and the yaw angle and pitch angle of the person are 5 degrees and 7 degrees respectively. The system can search for high-quality images that meet the following conditions: the capture time is between 14:02:05 and 14:02:25; the capture point is entrance A; the yaw angle of the person is between 5.1 degrees and 5.5 degrees, and the pitch angle is between 7.1 degrees and 7.5 degrees. If there are high-quality images that meet the conditions, all or part of them can be determined as the reference image of the current image to be repaired.

[0080] Through this embodiment, the reference image of each image to be repaired is selected according to the preset image conditions, improving the adaptability of the reference image of each image to be repaired.

[0081] In an exemplary embodiment, the image inpainting model includes: MAE-VQGAN (Masked Autoencoder-Vector Quantization Generative Adversarial Network). Here, MAE-VQGAN is a hybrid network for image inpainting. The masked autoencoder applies a mask to the defective area of the image to be inpainted during the encoding stage, and then learns and quantizes the image features. The vector quantization technique further discretizes these feature representations to enhance the model's expressive power and generalization ability. The core method of the image inpainting model (Impainting model) is MAE-VQGAN. Given an input image X ∈ R H*W*3 and a binary mask matrix m ∈ {0, 1} H*W , the goal of the image inpainting function f is to synthesize a new image y ∈ R H*W*3 , that is, fill in the masked positions of the mask matrix (y = f(x, m)). Among them, f is MAE-VQGAN.

[0082] MAE-VQGAN includes the MAE and VQGAN models. Its training process (the training process of the f function) can be combined Figure 9 . MAE-VQGAN will output the probability distribution p(z i |x i , m). Among them, the codebook of VQGAN is also used. The MAE-VQGAN network will output the token (mark) corresponding to each image patch of each image. The true token can map the picture to the corresponding visual token by using the encoder Encoder (per_trained) of VQGAN. The cross-entropy loss can be used as the loss function for training. The prediction of the visual token can be obtained through the argmax function: Then, map the token to pixels through the decoder of VQGAN to obtain the output picture. The decoder of VQGAN can be the VIM (Vision Transformer) decoder.

[0083] Optionally, input each image to be restored and the reference image corresponding to each image to be restored into an image restoration model to obtain the restored image corresponding to each image to be restored, including: respectively taking each image to be restored and the reference image corresponding to each image to be restored as the current image to be restored and the current reference image and inputting them into the image restoration model, so that the image restoration model performs the following image restoration operations on the current image to be restored and the current reference image: performing a Gaussian blur operation on the current reference image to obtain the current reference image after blur processing; using the reference image corresponding to each image to be restored after blur processing and the reference image corresponding to each image to be restored as input-output image examples, and using the masked autoencoder-vector quantization generative adversarial network to predict the value of each image position in the output image corresponding to the current image to be restored, and determining the predicted output image corresponding to the current image to be restored as the restored image corresponding to the current image to be restored.

[0084] In this embodiment, the blur processing of the reference image corresponding to each image to be restored can be performed by the image restoration model. For each image to be restored, it can be respectively used as the current image to be restored for image restoration, and the reference image corresponding to the current image to be restored is the current reference image. When performing image restoration on the current image to be restored, the current image to be restored and the current reference image can be input into the image restoration model. In the image restoration model, first, a Gaussian blur operation can be performed on the current reference image to obtain the current reference image after blur processing, and then, using the reference image corresponding to each image to be restored after blur processing and the reference image corresponding to each image to be restored as input-output image examples (Prompts), the masked autoencoder-vector quantization generative adversarial network predicts the value of each image position in the output image corresponding to the current image to be restored (that is, fills in the masked positions of the mask matrix), and the predicted output image corresponding to the current image to be restored is output as the restored image corresponding to the current image to be restored.

[0085] Through this embodiment, using the masked autoencoder-vector quantization generative adversarial network for image restoration can improve the convenience and quality of image restoration; the image restoration model performs a Gaussian blur operation on the reference image corresponding to the image to be restored, and then obtains the required image examples, which can improve the integration of the model.

[0086] In an exemplary embodiment, for each restored image corresponding to an image to be restored, it can be classified as captured data again. If the classification result is not an image to be clustered, image restoration can be performed again. In order to reduce the system clustering time and storage pressure while implementing the restoration image backflow clustering, a specified flag bit is used to mark the images that have been restored, and the restored images are verified (to check whether they meet the preset clustering conditions). If the verification passes, they participate in the subsequent portrait clustering process; otherwise, they can be directly recorded as waste film data.

[0087] Correspondingly, after inputting each image to be restored and the reference image corresponding to each image to be restored into the image restoration model, the above method further includes: adding a specified flag bit to the restored image corresponding to each image to be restored; verifying the restored image corresponding to each image to be restored to obtain the verification result of the restored image corresponding to each image to be restored; based on the specified flag bit, recording the image information of the restored images that do not meet the preset clustering conditions into the waste film table.

[0088] Here, the verification result of the restored image corresponding to each image to be restored is used to indicate whether the restored image corresponding to each image to be restored meets the preset clustering conditions (that is, whether portrait clustering is allowed). Based on the verification result of the restored image corresponding to each image to be restored, it is possible to determine the restored images that meet the preset clustering conditions and the restored images that do not meet the preset clustering conditions. Among them, for the restored images that meet the preset clustering conditions, they can be output to the portrait clustering system and participate in the portrait clustering process; for the restored images that do not meet the preset clustering conditions, their image information can be recorded into the waste film table.

[0089] Verifying the restored image corresponding to each image to be restored can be achieved by writing it into the captured data queue to participate in the classification process. Based on the specified flag bit, it is possible to distinguish between the restored images and the non-restored images. Correspondingly, based on the specified flag bit, it is possible to identify the restored images that do not meet the preset clustering conditions, so as to record the image information of the restored images that do not meet the preset clustering conditions into the waste film table.

[0090] For example, as Figure 10 shown, the restored waste film data can be input into the parsing module again as a captured image, and a flag bit P is given. After re-parsing, it enters the clustering system (that is, the portrait clustering system) again. At this time, there will be two results. One is that it passes the waste film verification module and enters the clustering stage, and the other is that it is still classified as waste film data. In this case, the data that is re-classified as waste film data and has the flag bit P is not restored again, but is directly sent to the waste film table and not clustered.

[0091] In this embodiment, by performing defective film verification on the repaired images and recording the image information of the repaired images that fail the defective film verification into the defective film table, the system clustering time and storage pressure can be reduced; by adding a flag bit to the repaired images and identifying the repaired images that fail the defective film verification based on the flag bit, the convenience of image processing can be improved.

[0092] In an exemplary embodiment, considering the obvious differences between human faces and human bodies, in order to improve the accuracy of portrait clustering, the face data and the human body data can be processed separately. Correspondingly, the captured data queue can include a face data queue and a human body data queue, and the classification processing is performed on the face data queue and the face data queue respectively. In order to reduce the requirements for the capture device, by analyzing the captured data, the face data and the human body data therein can be identified, the identified face data and human body data can be extracted, and then transmitted to the face data queue and the human body data queue respectively.

[0093] Correspondingly, before performing classification processing on the captured data in the captured data queue, the above method further includes: taking each captured data obtained from a group of capture devices as the current captured data and performing the following data processing operations: performing a parsing operation on the current captured data to obtain a data parsing result; in the case where the data parsing result indicates that the current captured data only contains face data, putting the face data in the current captured data into the face data queue; in the case where the data parsing result indicates that the current captured data only contains human body data, putting the human body data in the current captured data into the human body data queue; in the case where the data parsing result indicates that the current captured data contains both face data and human body data, putting the face data in the current captured data into the face data queue and putting the human body data in the current captured data into the human body data queue.

[0094] For each captured data, it can be used as the current captured data to perform a parsing operation to obtain a data parsing result. Here, the parsing operation is used to parse the face data and the human body data included in the current captured data. The parsing operation can be performed by a target recognition model, and the targets recognized by the target recognition model can be human faces and human bodies. The data parsing result can indicate at least one of the following: the current captured data only contains face data, the current captured data only contains human body data, the current captured data contains both face data and human body data, and the current captured data contains neither face data nor human body data.

[0095] For the case where the current captured data only contains face data, the face data in the current captured data can be put into the face data queue. The way to put the face data in the current captured data into the face data queue can be to directly put the current captured data into the face data queue, or to intercept the face part in the current captured data based on the recognized face detection frame and put the intercepted face data into the face data queue.

[0096] For the case where the current captured data only contains body data, the body data in the current captured data can be put into the body data queue. The way to put the body data in the current captured data into the body data queue can be to directly put the current captured data into the body data queue, or to intercept the body part in the current captured data based on the recognized body detection frame and put the intercepted body data into the body data queue.

[0097] For the case where the current captured data contains both face data and body data, the face data in the current captured data can be put into the face data queue, and the body data in the current captured data can be put into the body data queue. The way to put the face data in the current captured data into the face data queue and the body data in the current captured data into the body data queue can be: intercept the face part in the current captured data based on the recognized body detection frame and put the intercepted face data into the face data queue; intercept the body part in the current captured data based on the recognized body detection frame and put the intercepted body data into the body data queue.

[0098] For the case where the current captured data contains neither face data nor body data, the current captured data can be directly discarded. In addition, if the captured data of the same capture device within a period of time are all discarded, an alarm message can be sent to alarm the possible abnormality of the capture device.

[0099] For example, the face data and the body data can be transmitted to the face parsing operator and the body parsing operator respectively by the face message queue (i.e., the face data queue) and the body message queue (i.e., the body data queue) for analysis, and finally the corresponding feature values and various attribute values (such as face width, whether wearing a mask, etc.) are generated. The parsed face data and body data can be transmitted to different message queues downstream for subsequent quality verification and other processes.

[0100] Through this embodiment, by parsing the captured data of the capture device and storing the face data and the body data therein into the face data queue and the body data queue respectively based on the parsing result, the efficiency and accuracy of subsequent classification processing can be improved, and the setting requirements for the capture device can also be reduced.

[0101] In an exemplary embodiment, similar to the foregoing embodiments, the captured data queue includes a face data queue and a human body data queue. Among them, the classification process is performed on the face data queue and the face data queue respectively. To improve the convenience of data collection and the efficiency of processing captured data, different types of capture devices can be used to collect face data and human body data respectively. A set of capture devices includes a capture device for capturing face data and a capture device for capturing human body data. For example, the personnel capture data can be obtained by the point capture device and transmitted to the front-end storage. According to different types of capture devices, face data and human body data are collected respectively, and then transmitted to the face message queue and the human body message queue respectively.

[0102] Correspondingly, before classifying the captured data in the captured data queue, the above method further includes: putting the face data in the captured data obtained by the capture device for capturing face data in a set of capture devices into the face data queue; putting the human body data in the captured data obtained by the capture device for capturing human body data in a set of capture devices into the human body data queue.

[0103] For the captured data obtained from the capture device for capturing face data, it can be identified whether it contains face data, and the face data in the captured data obtained from the capture device for capturing face data is put into the face data queue. Similar to the foregoing embodiments, the captured data with recognized faces can be directly put into the face data queue, or the face part in the captured data can be intercepted based on the recognized face detection frame, and the intercepted face data is put into the face data queue.

[0104] For the captured data obtained from the capture device for capturing human body data, it can be identified whether it contains human body data, and the human body data in the captured data obtained from the capture device for capturing human body data is put into the human body data queue. Similar to the foregoing embodiments, the captured data with recognized human bodies can be directly put into the human body data queue, or the human body part in the captured data can be intercepted based on the recognized human body detection frame, and the intercepted human body data is put into the human body data queue.

[0105] Through this embodiment, using the capture device for capturing face data and the capture device for capturing human body data to collect face data and human body data respectively can improve the convenience of data collection and the efficiency of processing captured data.

[0106] In an exemplary embodiment, similar to the aforementioned embodiment, the captured data queue includes a face data queue and a body data queue, wherein the portrait clustering operation is performed on the face data queue and the body data queue respectively, and a group of portrait files obtained by performing the portrait clustering operation on the face data queue is a group of face files, and a group of portrait files obtained by performing the portrait clustering operation on the body data queue is a group of body files.

[0107] There is a certain association relationship between face files and body files. For example, face files and body files belong to the same person. In order to facilitate information processing (for example, information display), after performing portrait clustering operations on a group of images to be clustered and restored images that meet preset clustering conditions, the above method also includes: based on the specified association relationship between face data and body data, associating a group of face files and a group of body files.

[0108] Here, the specified association relationship is an association relationship specified based on at least one of the capture device, capture time and image features. For example, the specified association relationship may include: the corresponding capture devices are the same capture device, or the ratio of the device distance of the corresponding capture device to the time difference of the capture time meets the moving speed range of the human-shaped object; the distance of the position of the human-shaped object (which can be determined based on the image features) is less than or equal to the set distance threshold. Associating a group of face files and a group of body files may be: placing the face files and body files belonging to the same human-shaped object together.

[0109] For example, in the solution of adding person trajectories based on visual large model image restoration, in the clustering stage, such as Figure 11 As shown, the clustering process may include the following steps: obtaining a high-quality and effective face data stream to generate a face file; obtaining a high-quality and effective body data stream to generate a body file; merging face and body files. When generating a face file, high-quality face data output to the face message queue may be obtained, and the face data may be aggregated into a face file according to a clustering algorithm according to a certain time range (adjustable) and a similarity threshold; when generating a body file, high-quality body data output to the body message queue may be obtained, and the body data may be aggregated into a body file according to a clustering algorithm according to a certain time range (adjustable) and a similarity threshold; when merging a face and body file, the generated face file and body file may be associated according to the face-body association relationship analyzed from the image.

[0110] Through this embodiment, face data and body data are clustered respectively to obtain face files and body files, and the face files and body files are associated based on the association relationship between the face and the body, which can improve the accuracy of portrait clustering.

[0111] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0113] According to another aspect of the embodiments of this application, there is also provided a processing device for portrait files. This processing device for portrait files can be used to implement the method for processing portrait files provided in the above embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware, is also possible and contemplated.

[0114] Figure 12 is a structural block diagram of an optional processing device for portrait files according to the embodiments of this application. As Figure 12 shown in, this processing device for portrait files includes:

[0115] A classification unit 1202, configured to classify the captured data in the captured data queue to obtain a set of images to be clustered and a set of images to be repaired. Among them, the captured data queue is generated based on the captured data of a set of capture devices. A set of images to be clustered are captured data that meet the preset clustering conditions, and a set of images to be repaired are captured data that do not meet the preset clustering conditions. The preset clustering conditions are the conditions that the images allowed for portrait clustering meet;

[0116] A determination unit 1204, configured to determine a reference image corresponding to each to-be-restored image in a set of to-be-restored images from a set of to-be-clustered images based on a capture time and a capture location;

[0117] An input unit 1206, configured to input each to-be-restored image and the reference image corresponding to each to-be-restored image into an image restoration model to obtain a restored image corresponding to each to-be-restored image, where the image restoration model is a model for performing image restoration on an input to-be-restored image by using a blurred reference image and the reference image as input-output image examples;

[0118] A clustering unit 1208, configured to perform a portrait clustering operation on a set of to-be-clustered images and the restored images that meet a preset clustering condition to obtain a set of portrait files.

[0119] It should be noted that the classification unit 1202 in this embodiment may be configured to execute the above step S202, the determination unit 1204 in this embodiment may be configured to execute the above step S204, the input unit 1206 in this embodiment may be configured to execute the above step S206, and the clustering unit 1208 in this embodiment may be configured to execute the above step S208.

[0120] Through this application, the capture data in the capture data queue is classified to obtain a set of to-be-clustered images and a set of to-be-restored images, where the capture data queue is generated based on the capture data of a set of capture devices, the set of to-be-clustered images is the capture data that meets the preset clustering condition, the set of to-be-restored images is the capture data that does not meet the preset clustering condition, and the preset clustering condition is the condition satisfied by the images that allow portrait clustering; a reference image corresponding to each to-be-restored image in the set of to-be-restored images is determined from the set of to-be-clustered images based on the capture time and the capture location; each to-be-restored image and the reference image corresponding to each to-be-restored image are input into the image restoration model to obtain a restored image corresponding to each to-be-restored image, where the image restoration model is a model for performing image restoration on the input to-be-restored image by using the blurred reference image and the reference image as input-output image examples; a portrait clustering operation is performed on the set of to-be-clustered images and the restored images that meet the preset clustering condition to obtain a set of portrait files, which solves the problem of poor portrait aggregation effect caused by the loss of capture data in the portrait file processing method in the related art, improves the utilization rate of the capture data, and enhances the portrait aggregation effect.

[0121] In an exemplary embodiment, the classification unit includes: a first determination module configured to determine a parameter value corresponding to each capture data in the capture data queue and each evaluation parameter in a set of evaluation parameters, where the set of evaluation parameters includes an image quality parameter and at least one of the following: yaw angle, pitch angle; a second determination module configured to determine, as images to be clustered, the capture data in the capture data queue for which the parameter values corresponding to each evaluation parameter are all within the parameter value range of each evaluation parameter, thereby obtaining a set of images to be clustered; and a third determination module configured to determine, as images to be repaired, the capture data in the capture data queue for which the parameter value corresponding to the image quality parameter is not within the parameter value range of the image quality parameter and the parameter values corresponding to the other evaluation parameters in the set of evaluation parameters except the image quality parameter are all within the parameter value ranges of the other evaluation parameters, thereby obtaining a set of images to be repaired.

[0122] In an exemplary embodiment, the determination unit includes: a first execution module configured to perform the following processing operations on each image to be repaired respectively as the current image to be repaired to obtain a reference image corresponding to each image to be repaired: when there are images to be clustered that meet the preset image conditions in the set of images to be clustered, determining at least one image to be clustered that meets the preset image conditions as the reference image corresponding to the current image to be repaired, where the preset image conditions are that the time difference between the capture time and the capture time of the current image to be repaired is less than or equal to a preset time difference threshold, the capture point is the same as the capture point of the current image to be repaired, and the angular difference between the angular value of the specified angle and the angular value of the specified angle of the current image to be repaired is within a preset angular difference range, and the specified angle includes at least one of the following: yaw angle, pitch angle.

[0123] In an exemplary embodiment, the image repair model includes a masked autoencoder-vector quantization generative adversarial network, and the masked autoencoder-vector quantization generative adversarial network is a hybrid network for image repair. The input unit includes: a second execution module configured to input each image to be repaired and the reference image corresponding to each image to be repaired into the image repair model respectively as the current image to be repaired and the current reference image, so that the image repair model performs the following image repair operations on the current image to be repaired and the current reference image: performing a Gaussian blur operation on the current reference image to obtain the current reference image after blur processing; using the current reference image after blur processing corresponding to each image to be repaired and the reference image corresponding to each image to be repaired as input-output image examples, predicting the value of each image position in the output image corresponding to the current image to be repaired by the masked autoencoder-vector quantization generative adversarial network, and determining the predicted output image corresponding to the current image to be repaired as the repaired image corresponding to the current image to be repaired.

[0124] In an exemplary embodiment, the above-mentioned device further includes: an adding unit, configured to add a specified flag bit to the repaired image corresponding to each image to be repaired after inputting each image to be repaired and the reference image corresponding to each image to be repaired into the image repair model, where the specified flag bit is used to mark the image that has been repaired; a verification unit, configured to verify the repaired image corresponding to each image to be repaired to obtain the verification result of the repaired image corresponding to each image to be repaired, where the verification result of the repaired image corresponding to each image to be repaired is used to indicate whether the repaired image corresponding to each image to be repaired meets the preset clustering condition; a recording unit, configured to record the image information of the repaired image that does not meet the preset clustering condition into the waste film table based on the specified flag bit, where the waste film table is used to record the images that do not participate in the portrait clustering operation.

[0125] In an exemplary embodiment, the captured data queue includes a face data queue and a body data queue, where the classification process is performed on the face data queue and the face data queue respectively. The above-mentioned device further includes: an execution unit, configured to, before performing the classification process on the captured data in the captured data queue, use each captured data obtained from a group of capture devices as the current captured data to perform the following data processing operations: perform a parsing operation on the current captured data to obtain a data parsing result, where the parsing operation is used to parse the face data and body data included in the current captured data; in the case where the data parsing result indicates that the current captured data only includes face data, put the face data in the current captured data into the face data queue; in the case where the data parsing result indicates that the current captured data only includes body data, put the body data in the current captured data into the body data queue; in the case where the data parsing result indicates that the current captured data includes both face data and body data, put the face data in the current captured data into the face data queue, and put the body data in the current captured data into the body data queue.

[0126] In an exemplary embodiment, the captured data queue includes a face data queue and a body data queue, where the classification process is performed on the face data queue and the face data queue respectively; a group of capture devices includes a capture device for capturing face data and a capture device for capturing body data. The above-mentioned device further includes: a first putting unit, configured to, before performing the classification process on the captured data in the captured data queue, put the face data in the captured data obtained from the capture device for capturing face data in a group of capture devices into the face data queue; a second putting unit, configured to put the body data in the captured data obtained from the capture device for capturing body data in a group of capture devices into the body data queue.

[0127] In an exemplary embodiment, the captured data queue includes a face data queue and a human body data queue. Among them, the portrait clustering operation is performed separately for the face data queue and the human body data queue. A set of portrait files obtained by performing the portrait clustering operation on the face data queue is a set of face files, and a set of portrait files obtained by performing the portrait clustering operation on the human body data queue is a set of human body files. The above device further includes an association unit configured to, after performing the portrait clustering operation on a set of images to be clustered and the repaired images that meet the preset clustering conditions, associate a set of face files and a set of human body files based on a specified association relationship between the face data and the human body data, where the specified association relationship is an association relationship specified based on at least one of the capture device, the capture time, and the image features.

[0128] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: all the above modules are located in the same processor; or, the above-mentioned various modules are separately located in different processors in any combination form.

[0129] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, where the program, when running, executes the steps in any one of the above method embodiments.

[0130] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disc that can store computer programs.

[0131] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to execute the steps in any one of the above method embodiments through the computer program. In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0132] Specific examples in this embodiment may refer to the examples described in the above embodiments and the exemplary embodiments. This embodiment will not be elaborated here.

[0133] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes computer programs / instructions, and the computer programs / instructions contain program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1309 and / or installed from the removable medium 1311. When the computer program is executed by the central processing unit 1301, various functions provided by the embodiments of the present application are executed. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0134] Figure 13 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application. As Figure 13 shown, the computer system 1300 includes a CPU (Central Processing Unit) 1301, which can perform various appropriate actions and processes according to the programs stored in the ROM 1302 or the programs loaded from the storage part 1308 into the RAM 1303. In the random access memory 1303, various programs and data required for system operation are also stored. The central processing unit 1301, the read-only memory 1302, and the random access memory 1303 are connected to each other through a bus 1304. The I / O (Input / Output) interface 1305 is also connected to the bus 1304.

[0135] The following components are connected to the I / O interface 1305: an input part 1306 including a keyboard, a mouse, etc.; an output part 1307 including such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display), etc. and a speaker, etc.; a storage part 1308 including a hard disk, etc.; and a communication part 1309 including a network interface card such as a local area network card, a modem, etc. The communication part 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that the computer program read from it can be installed into the storage part 1308 as needed.

[0136] In particular, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1309, and / or installed from the removable medium 1311. When the computer program is executed by the central processing unit 1301, various functions defined in the system of the present application are executed.

[0137] It should be noted that Figure 13 The computer system 1300 of the electronic device shown is only an example, and should not bring any limitations to the functions and usage scope of the embodiments of the present application.

[0138] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to be implemented. In this way, the present application is not limited to any specific combination of hardware and software.

[0139] The above are only the preferred embodiments of the present application, and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for processing portrait files, characterized in that, Including: Classify the capture data in the capture data queue to obtain a set of images to be clustered and a set of images to be repaired. Among them, the capture data queue is generated based on the capture data of a set of capture devices. The set of images to be clustered are capture data that meet the preset clustering conditions, and the set of images to be repaired are capture data that do not meet the preset clustering conditions. The preset clustering conditions are the conditions that images allowing for portrait clustering satisfy; Based on the capture time and capture location, determine the reference image corresponding to each image to be repaired in the set of images to be repaired from the set of images to be clustered; Input each image to be repaired and the reference image corresponding to each image to be repaired into an image repair model to obtain the repaired image corresponding to each image to be repaired. Among them, the image repair model is a model for performing image repair on the input image to be repaired with the blurred reference image and the reference image as the input output image examples; Perform a portrait clustering operation on the set of images to be clustered and the repaired images that meet the preset clustering conditions to obtain a set of portrait files.

2. The method according to claim 1, wherein The classifying and processing the capture data in the capture data queue to obtain a set of images to be clustered and a set of images to be repaired includes: Determine the parameter value corresponding to each evaluation parameter in a set of evaluation parameters for each capture data in the capture data queue. Among them, the set of evaluation parameters includes an image quality parameter and at least one of the following: yaw angle, pitch angle; Determine the capture data in the capture data queue, for which the parameter values corresponding to each evaluation parameter are all within the parameter value range of each evaluation parameter, as the images to be clustered to obtain the set of images to be clustered; Determine the capture data in the capture data queue, for which the parameter value corresponding to the image quality parameter is not within the parameter value range of the image quality parameter and the parameter values corresponding to the other evaluation parameters in the set of evaluation parameters except the image quality parameter are all within the parameter value range of the other evaluation parameters, as the images to be repaired to obtain the set of images to be repaired.

3. The method according to claim 1, characterized in that, The determining the reference image corresponding to each image to be repaired in the set of images to be repaired from the set of images to be clustered based on the capture time and capture location includes: Take each image to be repaired as the current image to be repaired and perform the following processing operations to obtain the reference image corresponding to each image to be repaired: In the case where there are images to be clustered that meet the preset image conditions in the set of images to be clustered, determine at least one image to be clustered that meets the preset image conditions as the reference image corresponding to the current image to be repaired. Among them, the preset image conditions are that the time difference between the capture time and the capture time of the current image to be repaired is less than or equal to the preset time difference threshold, the capture location is the same as the capture location of the current image to be repaired, and the angle difference between the angle value of the specified angle and the angle value of the specified angle of the current image to be repaired is within the preset angle difference range. The specified angle includes at least one of the following: yaw angle, pitch angle.

4. The method according to claim 1, wherein The image inpainting model includes a masked autoencoder-vector quantization generative adversarial network, and the masked autoencoder-vector quantization generative adversarial network is a hybrid network for image inpainting; The step of inputting each image to be inpainted and the reference image corresponding to each image to be inpainted into the image inpainting model to obtain the inpainted image corresponding to each image to be inpainted includes: Respectively taking each image to be inpainted and the reference image corresponding to each image to be inpainted as the current image to be inpainted and the current reference image and inputting them into the image inpainting model, so that the image inpainting model performs the following image inpainting operations on the current image to be inpainted and the current reference image: Performing a Gaussian blur operation on the current reference image to obtain the current reference image after blur processing; Using the reference image corresponding to each image to be inpainted after blur processing and the reference image corresponding to each image to be inpainted as input-output image examples, and using the masked autoencoder-vector quantization generative adversarial network to predict the value of each image position in the output image corresponding to the current image to be inpainted, and determining the predicted output image corresponding to the current image to be inpainted as the inpainted image corresponding to the current image to be inpainted.

5. The method according to claim 1, characterized in that, After inputting each image to be inpainted and the reference image corresponding to each image to be inpainted into the image inpainting model, the method further includes: Adding a specified flag bit to the inpainted image corresponding to each image to be inpainted, where the specified flag bit is used to mark the image that has been inpainted; Verifying the inpainted image corresponding to each image to be inpainted to obtain the verification result of the inpainted image corresponding to each image to be inpainted, where the verification result of the inpainted image corresponding to each image to be inpainted is used to indicate whether the inpainted image corresponding to each image to be inpainted meets the preset clustering condition; Based on the specified flag bit, recording the image information of the inpainted image that does not meet the preset clustering condition into a waste film table, where the waste film table is used to record images that do not participate in the portrait clustering operation.

6. The method according to any one of claims 1 to 5, characterized in that The captured data queue includes a face data queue and a body data queue, and the classification processing is performed on the face data queue and the body data queue respectively; Before performing classification processing on the captured data in the captured data queue, the method further includes: Respectively taking each captured data obtained from the group of capture devices as the current captured data and performing the following data processing operations: Performing a parsing operation on the current captured data to obtain a data parsing result, where the parsing operation is used to parse the face data and body data included in the current captured data; When the data parsing result indicates that the current captured data only includes face data, putting the face data in the current captured data into the face data queue; When the data parsing result indicates that the current captured data only includes body data, putting the body data in the current captured data into the body data queue; When the data parsing result indicates that the current captured data includes face data and body data, put the face data in the current captured data into the face data queue, and put the body data in the current captured data into the body data queue.

7. The method according to any one of claims 1 to 5, characterized in that The captured data queue includes a face data queue and a body data queue. Among them, the classification process is performed on the face data queue and the face data queue respectively; the group of capture devices includes a capture device for capturing face data and a capture device for capturing body data; Before performing the classification process on the captured data in the captured data queue, the method further includes: Put the face data in the captured data obtained by the capture device for capturing face data from the group of capture devices into the face data queue; Put the body data in the captured data obtained by the capture device for capturing body data from the group of capture devices into the body data queue.

8. The method according to any one of claims 1 to 5, characterized in that, The captured data queue includes a face data queue and a body data queue. Among them, the portrait clustering operation is performed on the face data queue and the body data queue respectively. A group of portrait files obtained by performing the portrait clustering operation on the face data queue is a group of face files, and a group of portrait files obtained by performing the portrait clustering operation on the body data queue is a group of body files; After performing the portrait clustering operation on the group of images to be clustered and the repaired images that meet the preset clustering conditions, the method further includes: Based on the specified association relationship between the face data and the body data, associate the group of face files and the group of body files, where the specified association relationship is an association relationship specified based on at least one of the capture device, capture time, and image features.

9. A processing device for portrait files, characterized in that, Includes: A classification unit for classifying the captured data in the captured data queue to obtain a group of images to be clustered and a group of images to be repaired. Among them, the captured data queue is generated based on the captured data of a group of capture devices. The group of images to be clustered is the captured data that meets the preset clustering conditions, and the group of images to be repaired is the captured data that does not meet the preset clustering conditions. The preset clustering condition is the condition satisfied by the images that allow portrait clustering; A determination unit for determining, based on the capture time and capture location, the reference image corresponding to each image to be repaired in the group of images to be repaired from the group of images to be clustered; An input unit for inputting each image to be repaired and the reference image corresponding to each image to be repaired into an image repair model to obtain the repaired image corresponding to each image to be repaired. Among them, the image repair model is a model for performing image repair on the input image to be repaired with the blurred reference image and the reference image as the input output image example; A clustering unit for performing a portrait clustering operation on the group of images to be clustered and the repaired images that meet the preset clustering conditions to obtain a group of portrait files.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.