Picture archiving method, device, terminal equipment and computer readable storage medium
By screening facial images with qualified quality and combining target detection and bottom-up recognition methods to determine human features, the archiving reliability problem caused by changes in facial shooting angles and blurred features is solved, achieving higher image archiving accuracy.
Patent Information
- Application Number
- CN202111564466.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-12-20
AI Technical Summary
Existing image archiving methods have low reliability in calculating similarity when the face shooting angle changes or the facial features are blurred, resulting in poor archiving effect.
By screening facial images that meet the quality requirements, facial feature information is extracted, and the key points and posture information of the human head are determined by combining target detection and bottom-up recognition methods, and the facial and body feature information is archived.
Improves the reliability and accuracy of image archiving, ensuring the accuracy of facial and body feature recognition under different shooting angles and lighting conditions.
Smart Images

Figure CN114373203B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular relates to a method, apparatus, terminal device and computer-readable storage medium for archiving images. Background Art
[0002] Image archiving is the process of categorizing snapshot images captured by a camera into multiple archives. The goal of image archiving is to ensure that each archive contains the same subject. Image archiving is widely used in fields such as identity recognition and trajectory tracking. For example, when tracking multiple target individuals, multiple snapshot images can be archived, ensuring that each archive contains the same target individuals.
[0003] Existing image archiving methods typically extract facial features from captured images, calculate the similarity between two captured images based on this information, and then determine whether to group the two captured images into the same file based on the similarity. When the angle of the face is changed or the facial features are blurred, the calculated similarity becomes less reliable, resulting in poor image archiving results. Summary of the Invention
[0004] The embodiments of the present application provide a picture archiving method, apparatus, terminal device, and computer-readable storage medium, which can effectively improve the reliability of picture archiving results.
[0005] In a first aspect, an embodiment of the present application provides a method for archiving pictures, comprising: obtaining a plurality of pictures to be archived, each picture to be archived including a face and / or a body; for pictures to be archived including a face, detecting whether the face image included therein meets a predetermined quality condition; when a picture to be archived including a first face image meeting the quality condition is detected, extracting facial feature information of the first face image; for pictures to be archived including a body, identifying the head of the body through target detection processing to obtain a plurality of first head key point information; identifying the body by performing bottom-up recognition processing on the picture to be archived to obtain one or more second human posture information; determining human feature information of the image to be archived based on the first head key point information and the second human posture information; archiving the multiple images to be archived based on the facial feature information to obtain at least one facial file, and archiving the multiple images to be archived based on the human feature information to obtain at least one human file; matching the at least one facial file and the at least one human file based on the trajectory information to which the images to be archived belong, and merging the matched facial files and human files into one file, wherein the two images to be archived belonging to the same group of trajectory information contain the same photographed subject.
[0006] In each embodiment of the present application, the quality of the face image is screened before extracting facial feature information. Facial feature information is extracted only from face images that meet predetermined quality requirements. This allows for more accurate calculation of the facial similarity between two images and for archiving the images based on their facial feature information. Furthermore, by using different methods to determine the key head point information and body posture information of the human body image and then determining the human body feature information based on these two information, the recognition accuracy of the human body feature information is improved, thereby improving the reliability of image archiving.
[0007] In a second aspect, an embodiment of the present application provides a picture archiving device, comprising: a picture acquisition unit, for acquiring a plurality of pictures to be archived, each picture to be archived including a face and / or a human body; a quality detection unit, for detecting, for a picture to be archived including a face, whether the included face image meets a predetermined quality condition; a face feature recognition unit, for extracting face feature information of the first face image when a picture to be archived including a first face image meeting the quality condition is detected; a human body feature recognition unit, for: for a picture to be archived including a human body, identifying the head of the human body through target detection processing to obtain a plurality of first head key point information; and recognizing the human body by performing bottom-up recognition processing on the picture to be archived. The invention relates to a method for obtaining a first head key point information and a second human posture information by distinguishing the first head key point information and obtaining one or more second human posture information; determining the human feature information of the to-be-archived picture according to the first head key point information and the second human posture information; a first archiving unit, used to archive the multiple pictures to be archived according to the facial feature information to obtain at least one facial file, and to archive the multiple pictures to be archived according to the human feature information to obtain at least one human file; a second archiving unit, used to match the at least one facial file and the at least one human file according to the trajectory information to which the pictures to be archived belong, and merge the matched facial files and human files into one file, wherein the two pictures to be archived belonging to the same group of trajectory information contain the same shooting subject.
[0008] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a picture archiving method as described in any one of the first aspects above when executing the computer program.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and wherein the computer program, when executed by a processor, implements the image archiving method as described in any one of the above-mentioned first aspects.
[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the image archiving method described in any one of the above-mentioned first aspects.
[0011] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 This is a flow chart of the image archiving method provided in an embodiment of the present application;
[0014] Figure 2 This is a schematic diagram of key points of the human body when viewed from the front provided by an embodiment of the present application;
[0015] Figure 3 This is a schematic diagram of the key points of the human body corresponding to the left side provided in an embodiment of the present application;
[0016] Figure 4 is a schematic diagram of facial angles provided in an embodiment of the present application;
[0017] Figure 5 is a schematic diagram of the position of the shooting device provided in an embodiment of the present application;
[0018] Figure 6 This is a schematic diagram of the image archiving process provided by an embodiment of the present application;
[0019] Figure 7 This is a schematic diagram of the structure of a picture archiving device provided by an embodiment of the present application;
[0020] Figure 8 This is a structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0022] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0023] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting" depending on the context.
[0024] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0025] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.
[0026] See also Figure 1 , is a flow chart of a method for archiving pictures provided in an embodiment of the present application. As an example and not a limitation, the method may include the following steps:
[0027] S101: Acquire multiple pictures to be archived, each picture to be archived includes a face and / or a body.
[0028] In the embodiment of the present application, the multiple pictures to be archived can be taken by different shooting devices or by the same shooting device.
[0029] For example, in an application scenario, 10 cameras are installed in a shopping mall, and each camera takes 100 snapshots. These 10 cameras take a total of 1,000 snapshots. Some of these 1,000 snapshots include only facial parts, some include only human body parts, and some include the entire person (including both facial and human body parts). These 1,000 snapshots constitute a plurality of images to be archived. In subsequent embodiments of this application, images to be archived that include facial parts will be recorded as facial images, and images to be archived that include human body parts will be recorded as human body images.
[0030] In order to accurately extract feature information later, the picture including the entire person can also be segmented into a face picture including only the face part and a body picture including only the body part, and then feature extraction is performed on the segmented face picture and body picture respectively.
[0031] In practical applications, images to be archived can be manually sorted, such as by dividing them into face images and body images, or by manually dividing entire person images into face images and body images. Alternatively, trained recognition models can be used to sort images to be archived, without further limitation.
[0032] S102: For a picture to be archived that includes a face, detecting whether the included face image meets a predetermined quality condition.
[0033] The quality conditions are pre-set conditions used to preliminarily judge the quality of facial images and serve as the basis for the preliminary screening of facial images.
[0034] In one example, whether the facial image meets the quality condition can be determined based on the lighting conditions of the image. For example, detecting whether the included facial image meets the predetermined quality condition may include:
[0035] Obtain the light score of the face image, which is used to represent the brightness and darkness of the face image;
[0036] According to the light distribution of the face image, detect whether the face image meets the quality conditions.
[0037] Whether a facial image meets the quality condition can be detected through adaptive judgment. For example, the facial image's light score can be compared with a light score threshold to detect whether the facial image meets the quality condition. If the facial image's light score is greater than or equal to the light score threshold, the facial image meets the quality condition; if the facial image's light score is lower than the light score threshold, the facial image does not meet the quality condition. The light score threshold here can be a preset value, which is determined based on at least one of the parameters of the image acquisition device used to capture the image to be archived or the parameters of the shooting environment, and this embodiment of the present application is not limited to this.
[0038] S103 : When a picture to be archived including a first facial image meeting the quality condition is detected, facial feature information of the first facial image is extracted.
[0039] In one example, when a picture to be filed including a first facial image meeting a quality condition is detected, the detection of whether the facial image of the picture to be filed meets the quality condition is stopped, and instead facial feature information of the first facial image is extracted.
[0040] Predetermined quality conditions, such as light score thresholds, are used as a basis for determining whether to extract facial feature information from the first facial image. If the first facial image meets the predetermined quality conditions, it means that the overall quality of the first facial image is good and meets the quality requirements for facial recognition, and will not affect the reliability of subsequent image archiving. Therefore, the facial feature information of the first facial image can be extracted. The setting of the predetermined quality condition (such as the light score threshold) can be combined with actual conditions, based on experience or experimental data. The embodiment of the present application does not limit the setting method and basis of the quality condition.
[0041] If the first face image does not meet the predetermined quality condition, then the next picture to be archived that includes a face is checked to see if it meets the quality condition, or other processes are processed.
[0042] In each embodiment of the present application, the quality of the face image is screened before extracting facial feature information, and facial feature information is extracted only for face images that meet the quality requirements, so that the facial similarity between the two images can be calculated more accurately in the subsequent process to archive the images according to the facial feature information.
[0043] In this example, existing image feature extraction methods can be used to extract facial feature information. For example, feature information can be extracted using a trained neural network model, using a scale-invariant feature change method, or using a histogram of oriented gradients, etc., without specific limitations here.
[0044] It should be noted that when a picture to be archived includes both a face and a body, the picture to be archived is both a face picture and a body picture, and it is necessary to extract the face feature information and the body feature information of the picture to be archived respectively.
[0045] In one example, after extracting facial feature information for a face image to be archived, steps S102-S103 are repeated for the next face image to be archived until all face images to be archived are processed. Optionally, after processing a face image to be archived, other processes can be inserted, for example, a process for processing a human body image to be archived. After the other process is completed, the next face image to be archived is processed according to S102-S103.
[0046] S104 : For a picture to be archived that includes a human body, identify the human head through object detection processing to obtain a plurality of first head key point information.
[0047] Images to be archived that include human bodies are referred to as human body images. In the embodiment of the present application, in addition to archiving images based on faces, they are also archived based on human body feature information. Human body feature information is a human body feature that represents a person, such as human body posture.
[0048] The existing human posture recognition technology mainly includes two categories:
[0049] One type is a top-down recognition method, which first locates the approximate position of the human body and then identifies the specific posture. The most common method uses object detection to obtain the position frame of each person in the image. Based on the detected frame, the human skeleton key points are detected for each person, and the overall human posture is finally determined. Main methods include CPM, RMPE, Mask-RCNN, and GRMI.
[0050] The other type is a bottom-up recognition method. This method first finds all limbs and then groups them together. The main steps are to detect all key points in the image and then cluster them into different individuals through relevant strategies. Typical examples include: using heatmaps or regressing key point coordinates to calculate the information of each key point of the human posture; using Part Affinity Fields (PAF) to connect the calculated key points; and when multiple people are identified, using graph theory bipartite graph solution methods to obtain the human posture information of each person.
[0051] However, both of the above methods have the problem of low recognition accuracy. When performing human posture recognition, the embodiments of the present application use different methods to determine the head key point information and human posture information of the human body image and determine the human body feature information of the human body image based on the two, so that the recognition accuracy of the human body feature information is improved, thereby improving the reliability of image archiving.
[0052] In one example, identifying a human head through object detection processing to obtain a plurality of first head key point information includes:
[0053] S201, determining one or more head detection frames in the image to be archived through object detection processing; and
[0054] S202: Perform key point calculation processing on the first head detection frame to obtain a plurality of first head key point information.
[0055] In this example, target detection processing (for example, using Faster-RCNN or SSD) is used to obtain the first head detection frame of the human body in the image to be processed. The first head detection frame is a rectangular frame, and the head of the corresponding human body is inscribed in the rectangular frame as much as possible. When the image includes only one person, a first head detection frame is obtained through target detection processing; when the image includes multiple people, multiple first head detection frames are obtained through target detection processing. Target detection processing can be used to obtain relevant information about objects in the image, including the category and position of the object, such as whether it is a person or an object in the image, and the specific number, and the position information is usually represented by a bounding box (detection box). In this embodiment, target detection processing is used to obtain the bounding box of each human head.
[0056] Specifically, the target detection method that can be used in this embodiment is a two-stage detection method represented by FasterRCNN, R-FCN, etc. based on convolutional neural networks. This type of method mainly uses candidate windows plus deep learning classification, first extracting candidate areas, and then classifying the corresponding areas based on deep learning methods; end-to-end regression methods based on deep learning represented by YOLO, SSD, etc. can also be used. This type of method divides the image into several small squares in advance and performs feature extraction in the small squares. In addition, traditional detection methods can also be used in this embodiment to perform target detection processing, which are all within the scope of protection of the present invention.
[0057] After obtaining the first head detection frame, this example can obtain multiple first head key point information through key point calculation processing. The head key point information mainly refers to the position coordinate value of the pixel corresponding to the head key point (that is, the position coordinate value of the head key point) and the position coordinate value of each pixel in the pixel area corresponding to the head key point. The pixel area corresponding to the head key point refers to all the remaining pixels except the center pixel contained in the circular area with a radius of R, with the pixel corresponding to the head key point as the center pixel. The value of R is not limited. For example, R can take 3 times the standard deviation of the Gaussian function, or a certain proportion of the long side pixels of the image to be processed, such as 1 / 10 times.
[0058] In a specific example, the key point calculation process may include the following steps:
[0059] S301, calculating each key point information of the human body posture using a human body posture heat map or regressing key point coordinates;
[0060] S302, obtaining the head posture of the human body according to the calculated information of each key point;
[0061] S303, when the human head posture is frontal or back, the midpoints of each side of the first head detection frame are used as four head key points;
[0062] S304, when the head posture of the human body is the left side, the midpoint of the right longitudinal side, the lower left vertex, and the midpoint of the upper horizontal side of the first head detection frame are used as three head key points;
[0063] S305: When the head posture of the human body is right side profile, the midpoint of the left longitudinal side, the lower right vertex, and the midpoint of the upper horizontal side of the first head detection frame are used as three head key points.
[0064] First, step S301 is executed to calculate each key point information of the human body posture using the human body posture heat map or the regression key point coordinates. The use of the human body posture heat map to calculate each key point information of the human body posture and the use of the regression key point coordinates to calculate each key point information of the human body posture are well known to those skilled in the art and will not be described in detail here. In one example, 16 key point information is calculated (when the human body posture is frontal, the human body key points at this time are specifically referred to in Figure 2 ) or 15 key point information (when the human body posture is the left side, the key points of the head at this time are specifically referred to Figure 3 , where 1 is the top of the head, 2 is the left ear, and 3 is the chin). This is related to the computational model used, and the training of the computational model is related to manual labeling. The specific method of manual labeling will be introduced in detail later.
[0065] Then, step S302 is executed to obtain the head posture of the human body based on the information of each calculated key point. In this example, the calculated key points can be connected by PAF. When multiple people are identified, the bipartite graph solution method of graph theory (such as the Hungarian algorithm) is used to obtain the human posture information of each person; when only one person is identified, deep learning is used to obtain the human posture recognition information of each person based on the connected key points. PAF and the bipartite graph solution method of graph theory are well known to those skilled in the art and will not be described in detail here. Specifically, when step S301 determines that the head includes four key points, it indicates that the head posture of the human body is front or back; when step S301 determines that the head includes three key points and the key point located at the vertex of the head detection frame (i.e., the chin) is located at the leftmost side of the head detection frame, it indicates that the head posture of the human body is the left side; when step S301 determines that the head includes three key points and the key point located at the vertex of the head detection frame (i.e., the chin) is located at the rightmost side of the head detection frame, it indicates that the head posture of the human body is the right side.
[0066] At this point, the head pose of each person in the image to be processed can be determined, including front, back, left, or right profiles. Different head poses are then processed in different ways: when the head pose is front or back, the midpoints of each side of the first head detection frame are used as the four first head key points; when the head pose is left profile, the midpoint of the right longitudinal side of the first head detection frame, the lower left vertex, and the midpoint of the upper horizontal side are used as the three first head key points; when the head pose is right profile, the midpoint of the left longitudinal side of the first head detection frame, the lower right vertex, and the midpoint of the upper horizontal side are used as the three first head key points.
[0067] In another specific example, the key point calculation process may include the following steps:
[0068] S311, calculating each key point information of the human body posture using a human body posture heat map or regressing key point coordinates;
[0069] S312, obtaining the head posture of the human body based on the calculated information of each key point;
[0070] S313, when the head posture of the human body is sideways, performing horizontal expansion processing on the first head detection frame to obtain an expanded first head detection frame;
[0071] S314, when the head posture of the human body is frontal or back, the midpoints of each side of the first head detection frame are respectively used as four first head key points;
[0072] S315, when the head posture of the human body is the left side, the midpoint of the right longitudinal side, the lower left vertex, and the midpoint of the upper horizontal side of the expanded first head detection frame are used as three first preliminary head key point information;
[0073] S316, when the head posture of the human body is right side profile, the midpoint of the left longitudinal side, the lower right vertex, and the midpoint of the upper horizontal side of the expanded first head detection frame are used as three first preliminary head key point information;
[0074] S317 , when the head posture of the human body is the left profile or the right profile, performing a lateral contraction process corresponding to the lateral expansion process on the first preliminary head key point information to obtain first head key point information.
[0075] Compared with the previous example (S301-305), this example (S311-317) of key point calculation can refer to step S301, step S302, step S303, step S304 and step S305 respectively in steps S311, S312, S314, S315 and S316, except that step S313 and step S317 are added. When it is known that the head posture of the human body is the left side or the right side, step S313 is then executed to perform horizontal expansion processing on the first head detection frame, that is, to expand the first head detection frame in the horizontal direction of the face by a certain proportion (see Figure 3 As shown in the figure, the middle rectangular frame is the first head detection frame, and the expanded rectangular frame is the expanded first head detection frame). The value range of the expansion ratio can include 1.2-1.5. For example, while keeping the center point and the longitudinal length of the first head detection frame unchanged, the horizontal expansion is 1.2 times, 1.3 times, 1.4 times or 1.5 times, that is, the horizontal coordinate is expanded by 1.2 times, 1.3 times, 1.4 times or 1.5 times while taking the center of the first head detection frame as the origin and remaining unchanged, so that the expanded first head detection frame can completely cover the person's head, thereby improving the accuracy of subsequent recognition.
[0076] When step S315 or step S316 is executed, the three first preliminary head key point information corresponding to a human body can be obtained, and then step S317 is executed to perform a horizontal contraction processing on the three first preliminary head key point information obtained corresponding to the horizontal expansion processing in step S313, that is, the longitudinal coordinates remain unchanged, and the horizontal coordinates are contracted with the center of the first head detection frame as the origin. For example: when the expansion ratio of step S313 is 1.1 times, the contraction ratio is 1 / 1.1, and the three first preliminary head key point information after contraction are used as the three first head key point information.
[0077] When the head is in profile (not frontal or back), the first head detection frame is properly expanded and the resulting head key points are contracted to make the first head key point information more realistic, ultimately improving the accuracy of human pose recognition. It's important to reiterate that key point information includes the position coordinates of all pixels within a circular area with a radius of R, centered around the pixel corresponding to the key point.
[0078] So far, the first head key point information has been obtained.
[0079] S105 , identifying a human body by performing bottom-up recognition processing on the image to be archived to obtain one or more second human body posture information.
[0080] In this embodiment, the bottom-up identification process may include the following steps:
[0081] (1) Calculating the information of each key point of the human posture using a human posture heat map or regressing the key point coordinates. The specific implementation method can refer to step S301, and in the specific example, the information of step S301 can be directly obtained, so there is no need to repeat the execution;
[0082] (2) Partial affinity fields are used to connect the calculated key points. When multiple people are identified, the bipartite graph solution method of graph theory (such as the Hungarian algorithm) is used to obtain the human posture information of each person. The specific implementation method can refer to step S302, and in the specific example, the information of step S302 can be directly obtained, so there is no need to repeat it. Partial affinity fields are commonly used in human posture recognition technology. They are a non-parametric representation of the association of torso key points. It saves the position and direction information in the support area of the limb. The partial affinity field is a two-dimensional vector field for each limb: for each pixel in the area belonging to each limb (referring to the arm, leg, or torso), the two-dimensional vector encodes the direction from one part of the limb to another part. Each limb has a corresponding affinity field to connect its body parts. For details about partial affinity fields, please refer to the article "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields" (the URL of the article is: https: / / arxiv.org / abs / 1611.08050) or other public content, which will not be repeated here.
[0083] It should be noted that, in other embodiments of the present invention, other bottom-up recognition methods may be used to obtain the second human body posture information, which are all within the protection scope of the present invention.
[0084] At this point, the second human body posture information is obtained, and then S106 is executed.
[0085] S106: Determine human feature information of the image to be archived according to the first head key point information and the second human posture information.
[0086] In one example, determining the human feature information of the image to be archived based on the first head key point information and the second human posture information includes:
[0087] S401, extracting second head key point information and first torso key point information from second human body posture information;
[0088] S402, fusing the first head key point information and the second head key point information to obtain fused third head key point information; and
[0089] S403: Use the third head key point information and the first torso key point information as human body feature information of the image to be archived.
[0090] In this example, first, the second head key point information and the first torso key point information are extracted from the second human body posture information (S401), which is well known to those skilled in the art and will not be described in detail here. In this example, there are 12 first torso key points, and there are three key points on each limb. For details, please refer to Figure 2 The second head key points are 3 (when the human head posture is the left side or right side, please refer to Figure 3 ) or 4 (when the human head posture is front or back, please refer to Figure 2 ). It should be noted that, in this embodiment, it is necessary not only to extract the position coordinate values of the pixels corresponding to the head key points and the torso key points, but also to extract the position coordinate values of the pixels corresponding to the circular area with a radius of R centered on each head key point and each torso key point.
[0091] So far, the second head key point information has been obtained.
[0092] Next, the first head key point information and the second head key point information are fused to obtain fused third head key point information ( S402 ).
[0093] In this example, the Gaussian distribution value of each pixel in the entire image to be processed corresponding to each second head key point can be calculated first, and then the Gaussian distribution value of each pixel in the entire image to be processed corresponding to each fused third head key point can be calculated using the following formula. The calculated Gaussian distribution value is then converted into the position coordinate value of the pixel corresponding to each fused third head key point and the position coordinate value of the pixel corresponding to the circular area with a radius of R.
[0094] After solving the PAF connection problem using a bipartite graph solution based on graph theory, the connection vector field of all key points of the human posture is obtained. To improve positioning accuracy and further remove erroneous and redundant connections (these connections may be due to overlapping or hidden parts of the human body in the previous processing (e.g., step S105), making them relatively difficult to identify), the existing posture information is fused with the head position information of the bounding box to improve the accuracy of information positioning.
[0095] The fusion in this embodiment includes the following formula, that is, the Gaussian distribution value of each pixel in the entire image to be processed corresponding to each fused third head key point can be calculated by this formula:
[0096]
[0097]
[0098] wherein f k (x i ) is the Gaussian distribution value of the i-th pixel corresponding to the k-th third head key point after fusion, G is a bilinear interpolation function, R is the radius of the pixel region corresponding to the head key point, l k is the position coordinate value of the pixel corresponding to the k-th second head key point, x i is the Gaussian distribution value of the i-th pixel corresponding to the k-th second head key point, x j is the position coordinate value of the j-th pixel corresponding to the k-th first head key point, L k is the position coordinate value of the pixel corresponding to the k-th first head key point. Wherein the value range of i is 1-M, M is the number of pixels included in the circular region with radius R; the value range of j is 1-N, N is the total number of pixels in the image to be processed. In actual application, the value of R can be determined according to the Gaussian distribution value, and then the specific value of R in the previous step is determined.
[0099] In this example, when the head posture of the human body is front or back, the number of first head key points and second head key points is four, so that the value of k is 1, 2, 3 and 4. When the head posture of the human body is left side or right side, the number of first head key points and second head key points is three, so that the value of k is 1, 2 and 3.
[0100] By calculating three or four head key points, the third position coordinate value of the pixel corresponding to each head key point and the third position coordinate value of the pixel corresponding to each head key point, i.e. the third head key point information, are obtained.
[0101] Then, S403 is executed, and the third head key point information and the first torso key point information are taken as the human body feature information of the picture to be archived.
[0102] In S403, the three or four third head key point information obtained in S402 and the twelve first torso key point information obtained in S105 are taken as the human body feature information of the picture to be archived.
[0103] In S104-S106, the head key point information and the human body posture information of the human body picture are determined by different methods, and the human body feature information of the human body picture is determined based on the two, so that the recognition accuracy of the human body feature information is improved, thereby improving the reliability of the picture archiving. In addition, in the embodiments of S104-S106, when the human body posture recognition is performed, the first head key point information of the to-be-processed image is obtained through the target detection processing and the key point calculation processing, the second head key point information and the first torso key point information of the to-be-processed image are obtained through the bottom-up recognition processing, then the first head key point information and the second head key point information are fused to obtain the third head key point information, and finally the third head key point information and the first torso key point information are taken as the human body posture recognition result. As can be seen, the embodiments effectively combine the data labeling method and the human body posture prediction algorithm, and use the facial features which are easy to extract as the main target detection positioning features, thereby improving the recognition accuracy of the model.
[0104] Up to now, the face and human body feature extraction is realized for the to-be-archived face picture and the to-be-archived human body picture.
[0105] In the above embodiments, the to-be-archived pictures are divided into to-be-archived face pictures and to-be-archived human body pictures to perform steps S102-S106, but it can be understood that the pre-classification can not be performed, but the detection can be performed in the processing process of steps S102-S106. For example, after step S101, for each to-be-archived picture, first, it is detected whether the to-be-archived picture contains a face and / or a human body, if the face is contained, step S102 is entered, and if the human body is contained, step S104 is entered. Similarly, for the picture containing both the face and the human body, the processing of steps S102-S103 and the processing of steps S104-S106 are performed on it, and the order of performing the two parts of processing is not limited.
[0106] S107, archiving the plurality of to-be-archived pictures according to the face feature information to obtain at least one face archive, and archiving the plurality of to-be-archived pictures according to the human body feature information to obtain at least one human body archive.
[0107] In S107, the face pictures and the human body pictures in the to-be-archived pictures are respectively archived. In an embodiment, the step of archiving the face pictures comprises:
[0108] Calculating the face feature similarity between the face feature information of each two face pictures; if the face feature similarity is greater than a first preset threshold, the two face pictures are summarized into the same face archive.
[0109] The distance between the facial feature information of two facial images can be calculated using distance calculation methods such as Euclidean distance and Mahalanobis distance, and then the distance can be subtracted from a preset value to obtain the facial feature similarity. The facial feature similarity between the facial feature information of two facial images can also be calculated using similarity calculation methods such as cosine similarity, Pierce correlation coefficient, and Jaccard similarity coefficient. This application does not specifically limit the method for calculating facial feature similarity.
[0110] In practical applications, it's common to capture a face in profile. Because profiles contain less facial feature information than full-face images, the facial feature similarity between two profile images may be greater than the facial feature similarity between two full-face images. In this case, the calculated facial feature similarity is inaccurate.
[0111] To solve the above problem, in one embodiment, the step of archiving facial images includes:
[0112] Calculate the image quality value of each of the multiple face images; calculate the face fusion similarity between every two face images based on the image quality value and face feature information; archive the multiple face images based on the face fusion similarity to obtain at least one face file.
[0113] Face images typically exist in a variety of states, including angle, size, presence of a mask, and image clarity. Each state of a face image is used as a quality parameter for evaluating image quality. The values of each quality parameter are counted, and then the image quality value of the face image is calculated based on the parameter values.
[0114] For example, see angle Figure 4 , is a schematic diagram of the face angle provided by the embodiment of the present application. Figure 4 As shown in the figure, face angles include roll, pitch, and yaw. Roll represents the angle of rotation along the head's front-to-back axis, pitch represents the angle of rotation along the head's left-to-right axis, and yaw represents the angle of rotation along the head's bottom-to-top axis. These three angles can represent the rotation angle of the face relative to the camera.
[0115] Optionally, the image quality value may be calculated by:
[0116] For each face image, obtain parameter values of multiple image quality parameters of the face image; and perform weighted summation of the parameter values to obtain an image quality value of the face image.
[0117] Specifically, the formula Calculate the image quality value. Where Q represents the image quality value, n is the number of quality parameters, x iis the parameter value of the i-th quality parameter, ω i is the weight corresponding to the parameter value of the i-th quality parameter.
[0118] With the above Figure 4 Taking the angle of the embodiment as an example, the three angle values of roll, pitch and yaw can be used as a parameter value respectively. For the size of the picture, the length and width of the picture can be used as a parameter value respectively; the area of the picture can also be used as a parameter value. For whether to wear a mask, parameter values can be set separately for the cases of wearing a mask and not wearing a mask. For example, the parameter value for the case of wearing a mask is set to 1, and the parameter value for the case of not wearing a mask is set to 0. For the clarity of the picture, the resolution of the picture can be used as a parameter value. It should be noted that in addition to the several quality parameters listed in the embodiments of the present application, other quality parameters that affect the quality of the picture can also be selected, which are not specifically limited here. The angle value of the face picture can be obtained by identifying the existing face angle recognition model, which is not specifically limited here. The situation of a face wearing a mask can be obtained by identifying the existing face mask recognition model, which is not specifically limited here.
[0119] The greater the impact of a quality parameter on image quality, the greater its corresponding weight. For example, if the angle of a face in an image has a greater impact on image quality, while the size of the image has a smaller impact, then the weight corresponding to the angle should be increased, while the weight corresponding to the size should be decreased.
[0120] To improve the accuracy of image quality calculation, the weights can be continuously learned. For example, the weights can be learned using feature similarities between images and parameter values of multiple quality parameters for each image.
[0121] Through the above method, the image quality factor of the face image is taken into account in the calculation of the facial feature similarity, which effectively improves the accuracy of the facial similarity.
[0122] Based on the above description of image quality values, optionally, the calculation method of face fusion similarity may include: weighted summing the facial feature similarity between each two face images and the image quality value of each two face images to obtain the face fusion similarity between each two face images.
[0123] Optionally, another method for calculating face fusion similarity includes:
[0124] The facial feature similarity between each two face images is calculated based on the facial feature information; multiple face images are divided into multiple image groups based on the image quality values; a preset coefficient matrix is obtained, the coefficient matrix includes weight coefficients between the image groups to which each two face images belong; and the facial fusion similarity between each two face images is calculated based on the facial feature similarity and the weight coefficient between each two face images.
[0125] The image groups may be divided in a manner as follows: presetting a division range of image quality values; and dividing face images corresponding to image quality values within the division range into one image group.
[0126] For example, assume that the image quality values are divided into ranges of 0-50, 50-80, and 80-100. The image quality values of face images A, B, C, and D are 30, 60, 70, and 90, respectively. Face image A belongs to the first image group, face images B and C belong to the second image group, and face image D belongs to the third image group. It should be noted that the above is only an example of image group division and does not specifically limit the range of image quality values.
[0127] In the embodiment of the present application, the coefficient matrix can be pre-set manually, or calculated based on actual experience, and can also be continuously adjusted during the actual application process.
[0128] In an embodiment of the present application, the face similarity and the weight coefficient can be multiplied to obtain the face fusion similarity. For example, assume that there are three face pictures A, B, and C, and the image quality values of A, B, and C are 40, 70, and 90, respectively. The three face pictures are divided into two picture groups according to the image quality value. Specifically, the face pictures with an image quality value greater than 50 are divided into one picture group (high-quality picture group), and the face pictures with an image quality value less than 50 are divided into one picture group (low-quality picture group). That is, A belongs to the low-quality picture group, and B and C belong to the high-quality picture group. In the preset coefficient matrix, the weight coefficient between the high-quality picture group and the low-quality picture group is 0.9, the weight coefficient between the high-quality picture group and the high-quality picture group is 0.4, and the weight coefficient between the low-quality picture group and the low-quality picture group is 0.5.
[0129] Calculate the face fusion similarity between A and B: calculate the facial feature similarity between A and B; the weight coefficient between the low-quality image group to which A belongs and the high-quality image group to which B belongs is 0.9; multiply the facial feature similarity between A and B by 0.9 to obtain the face fusion similarity between A and B.
[0130] Calculate the face fusion similarity between A and C: Calculate the facial feature similarity between A and B; the weight coefficient between the low-quality image group to which A belongs and the high-quality image group to which C belongs is 0.9; multiply the facial feature similarity between A and C by 0.9 to obtain the face fusion similarity between A and C.
[0131] Calculate the face fusion similarity between B and C: calculate the facial feature similarity between B and C; the weight coefficient between the high-quality image group to which B belongs and the high-quality image group to which C belongs is 0.4; multiply the facial feature similarity between B and C by 0.4 to obtain the face fusion similarity between B and C.
[0132] As can be seen from the above example, by modifying the facial feature similarity within each image quality value range, the facial feature similarity measurement for each quality range is unified. This allows facial images to be archived using a unified threshold, avoiding unreasonable archiving results caused by inconsistent feature similarity measurements.
[0133] In one embodiment, the step of archiving a human body picture includes:
[0134] Calculate the human feature similarity between the human feature information of each two human body pictures; if the human feature similarity is greater than a second preset threshold, the two human body pictures are classified into the same human body file.
[0135] In actual applications, there may be two people wearing similar clothes, which will result in a high similarity of human features between the two human body images, resulting in inaccurate final archiving results.
[0136] To solve the above problem, in one embodiment, the step of archiving human body pictures includes:
[0137] Calculate the human feature similarity between every two human body pictures based on the human feature information; calculate the spatiotemporal similarity between every two human body pictures; calculate the human body fusion similarity between every two human body pictures based on the human feature similarity and the spatiotemporal similarity; archive multiple human body pictures based on the human body fusion similarity to obtain at least one human body file.
[0138] The spatiotemporal similarity includes similarity in temporal information and similarity in spatial information. In the embodiments of the present application, time may refer to the time it takes for a camera to capture a target object and obtain a human body image. Because multiple cameras may be present in a real scene, each installed in a different position, there may be an actual distance between them. This actual distance constitutes the spatial information.
[0139] For example, see Figure 5, is a schematic diagram of the position of the shooting device provided in the embodiment of the present application. Figure 5 As shown, the position points of camera A and camera B are obtained, where the position point of camera A is the intersection point O1 of the center of the field of view of camera A and the center line of the illuminated road, and the position point of camera B is the intersection point O2 of the center of the field of view of camera B and the center of the illuminated road. The actual distance between cameras A and B is the actual distance from O1 to O2 (as shown in Figure 2). Figure 5 The line segments O1M, MN and NO2 are shown).
[0140] Optionally, the spatiotemporal similarity can be calculated by:
[0141] The actual distance between the shooting devices corresponding to each two human body pictures is calculated; the time similarity between the shooting times of each two human body pictures is calculated; and the spatiotemporal similarity between each two human body pictures is calculated based on the actual distance and the time similarity.
[0142] The method for calculating the actual distance may be: determining the position points of the two shooting devices in the application scene map, determining the path between the two position points in the application scene map, calculating the actual length of the path, and determining the actual length as the actual distance.
[0143] Temporal similarity can be calculated by multiplying the shooting times of two human images to obtain the temporal similarity between them. Alternatively, temporal similarity can be calculated by calculating the similarity between the shooting times of two human images. For example, the difference between the two shooting times can be calculated and used as the temporal similarity. Alternatively, the cosine similarity between the two shooting times can be calculated and used as the temporal similarity. Of course, other similarity calculation methods, such as Euclidean distance and Mahalanobis distance, can also be used, and are not specifically limited here.
[0144] Optionally, one implementation method for calculating the spatiotemporal similarity between two human body images based on the actual distance and time similarity may be:
[0145] The calculated actual distance is normalized to reduce the distance to between 0 and 1, and then the normalized distance is subtracted from 1 to obtain the spatial similarity; then the spatial similarity and temporal similarity between each two human body images are multiplied to obtain the spatiotemporal similarity between each two human body images.
[0146] One implementation method of calculating the human body fusion similarity between every two human body pictures based on the human body feature similarity and the spatiotemporal similarity is: multiplying the human body feature similarity and the spatiotemporal similarity between every two human body pictures to obtain the human body fusion similarity between every two human body pictures.
[0147] Further, the plurality of human body pictures are archived according to the human body fusion similarity, to obtain at least one human body archive, including: if the human body fusion similarity between two human body pictures is greater than a second preset threshold, the two human body pictures are classified into the same human body archive.
[0148] In S107, a community map cutting method (such as infomap, louvain algorithm, etc.) can also be used for post-processing to improve the accuracy of the archive.
[0149] In S108, the at least one face archive and the at least one human body archive are matched according to the trajectory information to which the to-be-archived picture belongs, and the matched face archive and human body archive are merged into one archive, wherein the two to-be-archived pictures belonging to the same group of trajectory information contain the same shooting object.
[0150] In one embodiment, the implementation of the matching of the face archive and the human body archive includes:
[0151] For any human body archive, the matching value between the human body archive and each face archive is calculated respectively, and the face archive and the human body archive corresponding to the maximum matching value are merged into one archive.
[0152] The matching value represents the number of face pictures in the human body archive that belong to the trajectory information corresponding to the face archive.
[0153] Optionally, the face archive and the human body archive corresponding to the larger N matching values can also be merged into one archive. In actual application, the size of N value can be determined according to the accuracy of the archive. The larger the N value, the lower the archive accuracy; the smaller the N value, the higher the archive accuracy.
[0154] In the embodiment of the application, when the human body archive and the face archive are matched, the trajectory information can be used. The plurality of snapshot pictures of the same shooting object obtained by a single shooting device belong to the same group of trajectory information. The trajectory information can be obtained by using the existing trajectory tracking technology, which is not limited here.
[0155] In actual application, a shooting device can track and shoot a certain shooting object by using a tracking algorithm. In the tracking and shooting process, a plurality of snapshot pictures are obtained, which constitute a group of trajectory information. The plurality of snapshot pictures of a group of trajectory information can include face pictures and human body pictures. The face pictures and the human body pictures belonging to the same group of trajectory information are matched pictures.
[0156] Exemplarily, the human body archive A is matched with each face archive. It is assumed that there are two face archives B and C in total, wherein the human body archive A includes 10 human body pictures, and each of the face archives B and C includes 10 face pictures. 8 human body pictures in the human body archive A are matched with 8 face pictures in the face archive B, that is, 8 human body pictures in A belong to the track information corresponding to B, and the matching value between A and B is 8; 3 human body pictures in the human body archive A are matched with 3 face pictures in the face archive C, that is, 3 human body pictures in A belong to the track information corresponding to C, and the matching value between A and C is 3. The face archive corresponding to the maximum matching value (that is, 8) and the human body archive A are merged into one archive, that is, the face archive B corresponding to the matching value 8 and the human body archive A are merged into one archive.
[0157] In the above embodiment, the pictures to be archived are divided into face pictures to be archived and human body pictures to be archived to perform steps S102-S106, but it can be understood that the detection can also be performed in the processing of steps S102-S106 without pre-classification. For example, after step S101, for each picture to be archived, it is first detected whether the picture to be archived contains a face and / or a human body, if it contains a face, it enters step S102, and if it contains a human body, it enters step S104. It should be noted that for pictures containing both faces and human bodies, both steps S102-S103 and steps S104-S106 are performed on the pictures, and the order of performing the two parts of processing is not limited.
[0158] Referring to Figure 6 , it is a schematic diagram of a picture archiving process provided by the embodiment of the present application. As Figure 6 indicated, the face fusion similarity is obtained according to the picture quality and the face picture feature (that is, face feature information), the face pictures are archived according to the face fusion similarity, and the face archive is obtained. The human body fusion similarity is obtained according to the spatiotemporal similarity and the human body picture feature (that is, human body feature information), the human body pictures are archived according to the human body fusion similarity, and the human body archive is obtained. Finally, the human body archive and the face archive are aggregated according to the track information.
[0159] Through the above method, the human body feature information is considered on the basis of the face feature information; and the picture quality is considered when archiving the face, and the spatiotemporal information is considered when archiving the human body, the dimension of the feature information is increased, the inaccuracy of the similarity caused by the inaccuracy of the single-dimensional feature information is avoided, and the reliability of the picture archiving result is effectively improved.
[0160] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0161] Corresponding to the picture archiving method described in the above embodiment, Figure 7 This is a structural block diagram of the picture archiving device provided in an embodiment of the present application. For the sake of convenience, only the parts related to the embodiment of the present application are shown.
[0162] Reference Figure 7 , the device comprises:
[0163] The image acquisition unit 71 is used to acquire multiple images to be archived, each of which includes a face and / or a body;
[0164] The quality detection unit 72 is used to detect whether the face image to be archived meets the predetermined quality conditions;
[0165] The facial feature recognition unit 73 is configured to extract facial feature information of the first facial image when detecting a picture to be archived that includes the first facial image meeting the quality condition;
[0166] The human feature recognition unit 74 is configured to: for a picture to be filed that includes a human body, identify the human head through object detection processing to obtain a plurality of first head key point information; identify the human body through bottom-up recognition processing on the picture to be filed to obtain one or more second human posture information; and determine human feature information of the picture to be filed based on the first head key point information and the second human posture information;
[0167] a first archiving unit 75 configured to archive the plurality of images to be archived based on the facial feature information to obtain at least one facial archive, and to archive the plurality of images to be archived based on the body feature information to obtain at least one body archive;
[0168] The second archiving unit 76 is used to match the at least one face file and the at least one body file according to the trajectory information of the pictures to be archived, and merge the matched face files and body files into one file, wherein the two pictures to be archived belonging to the same set of trajectory information contain the same subject.
[0169] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0170] In addition, Figure 7 The picture archiving apparatus shown can be a software unit, a hardware unit, or a software and hardware combined unit built into an existing terminal device, can be integrated into the terminal device as an independent plug-in, or can exist as an independent terminal device.
[0171] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the above described functions. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0172] Figure 8 is a structural schematic diagram of a terminal device provided by the present application. As Figure 8 shown, the terminal device 8 of this embodiment includes at least one processor 80 Figure 8 only one processor is shown), a memory 81, and a computer program 82 stored in the memory 81 and executable on the at least one processor 80, wherein the processor 80 executes the computer program 82 to implement the steps in any of the above picture archiving method embodiments.
[0173] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and other computing devices. The terminal device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 8 is only an example of the terminal device 8 and does not constitute a limitation on the terminal device 8, and can include more or fewer components than shown, or combine certain components, or different components, for example, can also include input and output devices, network access devices, etc.
[0174] The processor 80 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0175] The memory 81 can be an internal storage unit of the terminal device 8 in some embodiments, for example, a hard disk or a memory of the terminal device 8. The memory 81 can also be an external storage device of the terminal device 8 in other embodiments, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 81 can include both an internal storage unit and an external storage device of the terminal device 8. The memory 81 is used to store an operating system, application programs, a boot loader, data, and other programs, for example, program codes of the computer program, etc. The memory 81 can also be used to temporarily store data that has been output or will be output.
[0176] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above-mentioned various method embodiments.
[0177] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is enabled to implement the steps in the above-mentioned various method embodiments.
[0178] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electric carrier signal or a telecommunication signal.
[0179] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0180] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0181] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0182] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0183] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for archiving pictures, characterized in that: include: Acquire multiple images to be archived, each of which includes a face and / or a body; For pictures to be archived that include human faces, detecting whether the included human face images meet predetermined quality conditions; When detecting a picture to be archived that includes a first facial image that meets the quality condition, extracting facial feature information of the first facial image; For a picture to be archived that includes a human body, identifying the head of the human body through object detection processing to obtain a plurality of first head key point information; Identifying the human body by performing bottom-up recognition processing on the image to be archived to obtain one or more second human body posture information; Determining human feature information of the image to be archived according to the first head key point information and the second human posture information; Archiving the plurality of images to be archived according to the facial feature information to obtain at least one facial file, which includes: Calculating an image quality value of each of a plurality of face pictures, wherein the face pictures are pictures to be archived that include faces; Calculating the face fusion similarity between every two face images according to the image quality value and the facial feature information; Archiving the pictures to be archived that include faces among the plurality of pictures to be archived according to the face fusion similarity to obtain the at least one face file; Archiving the plurality of images to be archived according to the human body feature information to obtain at least one human body archive; and matching the at least one face file and the at least one body file according to the trajectory information of the pictures to be archived, and merging the matched face files and body files into one file, wherein the two pictures to be archived belonging to the same set of trajectory information contain the same subject; The calculating the face fusion similarity between each two pictures to be archived that include faces based on the image quality value and the face feature information includes: Calculating the facial feature similarity between every two pictures to be archived that include faces based on the facial feature information; dividing the plurality of face images into a plurality of image groups according to the image quality values; Obtaining a preset coefficient matrix, where the coefficient matrix includes weight coefficients between the picture groups to which each two pictures to be archived belong; The face fusion similarity between each two pictures to be archived is calculated according to the facial feature similarity between each two pictures to be archived and the weight coefficient.
2. The image archiving method according to claim 1, characterized in that: The identifying the head of the human body through target detection processing to obtain a plurality of first head key point information includes: Determining one or more head detection frames in the image to be archived through object detection processing; and Key point calculation processing is performed on the head detection frame to obtain the plurality of first head key point information.
3. The image archiving method according to claim 1, characterized in that: The determining of the human feature information of the image to be archived according to the first head key point information and the second human posture information includes: Extracting second head key point information and first torso key point information from the second human body posture information; fusing the first head key point information and the second head key point information to obtain fused third head key point information; and The third head key point information and the first torso key point information are used as human feature information of the image to be archived.
4. The image archiving method according to claim 2, wherein: The key point calculation process includes: Use human body posture heat map or regression key point coordinates to calculate each key point information of human body posture; Obtain the head posture of the human body based on the information of each calculated key point; When the head posture of the human body is frontal or back, the midpoints of each side of the head detection frame are used as four head key points; When the head posture of the human body is the left side, the midpoint of the right longitudinal side of the head detection frame, the lower left vertex and the midpoint of the upper horizontal side are used as the three head key points; When the head posture of the human body is the right side, the midpoint of the left longitudinal side of the head detection frame, the lower right vertex and the midpoint of the upper horizontal side are used as the three head key points.
5. The image archiving method according to claim 1, wherein: Archiving the plurality of images to be archived according to the human body feature information to obtain at least one human body archive includes: Calculating the human body feature similarity between every two human body pictures according to the human body feature information, wherein the human body pictures are pictures to be archived that include human bodies; Calculating the spatiotemporal similarity between every two human body pictures; Calculating the human body fusion similarity between every two human body pictures according to the human body feature similarity and the spatiotemporal similarity; The human body pictures in the plurality of pictures to be archived are archived according to the human body fusion similarity to obtain the at least one human body file.
6. A picture archiving device, characterized in that: include: A picture acquisition unit, configured to acquire a plurality of pictures to be archived, each picture to be archived including a face and / or a body; A quality detection unit, configured to detect whether the face image to be archived, including the face, meets a predetermined quality condition; a facial feature recognition unit, configured to extract facial feature information of the first facial image when detecting a picture to be archived that includes the first facial image meeting the quality condition; A human feature recognition unit, configured to identify a head of a person in a picture to be archived by performing target detection processing to obtain a plurality of first head key point information; Identifying the human body by performing bottom-up recognition processing on the image to be archived to obtain one or more second human body posture information; determining human body feature information of the image to be archived based on the first head key point information and the second human body posture information; A first archiving unit is configured to archive the plurality of images to be archived based on the facial feature information to obtain at least one facial archive, comprising: calculating an image quality value of each of the plurality of facial images, wherein the facial images are images to be archived that include faces; Calculating the face fusion similarity between every two face images according to the image quality value and the facial feature information; Archiving the pictures to be archived that include faces among the plurality of pictures to be archived according to the face fusion similarity to obtain the at least one face file; The calculating the face fusion similarity between each two pictures to be archived that include faces based on the image quality value and the face feature information includes: Calculating the facial feature similarity between every two pictures to be archived that include faces based on the facial feature information; dividing the plurality of face images into a plurality of image groups according to the image quality values; Obtaining a preset coefficient matrix, where the coefficient matrix includes weight coefficients between the picture groups to which each two pictures to be archived belong; Calculating the face fusion similarity between each two pictures to be archived according to the face feature similarity between each two pictures to be archived and the weight coefficient; The second archiving unit is configured to match the at least one face file and the at least one body file according to the trajectory information to which the pictures to be archived belong, and merge the matched face files and body files into one file, wherein the two pictures to be archived belonging to the same set of trajectory information contain the same subject.
7. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Archiving method and device
CN109800674A
Method and device for positioning tracked target
CN111476820A
Human body posture recognition method and device
CN111797791A