Face recognition method and system based on face multi-label algorithm
By constructing a face multi-label algorithm, using sliding window scanning and multi-layer feature resampling technology, combined with knowledge graph matching, the problem of insufficient static attributes and dynamic motion recognition of face recognition in complex scenarios is solved, and efficient and accurate face recognition is achieved.
Patent Information
- Application Number
- CN202510493296.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art face recognition system in complex scenarios is difficult to accurately recognize the static properties and dynamic actions of the face at the same time, resulting in insufficient recognition accuracy and traditional methods are inefficient in complex contexts.
Using a face multi-label algorithm, the key points of the face are obtained through preliminary scanning of the sliding window, a global coverage path is constructed, and a multi-layer feature resampling and knowledge graph matching is used to obtain action tags and static attribute tags to achieve comprehensive description and optimized recognition.
It improves the accuracy and efficiency of face recognition, can accurately capture the overall and detailed features of the face in complex scenarios, and enhances the adaptability of the system and the reliability of recognition.
Smart Images

Figure CN120279588A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and in particular, to a face recognition method and system based on a face multi-label algorithm. Background Art
[0002] With the improvement of urban security requirements, security monitoring systems have been widely deployed. In complex scenarios, it is not only necessary to identify the face identity of a person, but also necessary to know information such as the person's behavior actions and related attributes to quickly determine whether there are potential security hazards. In practical applications, the face image data collected has a high degree of complexity and diversity. The expression changes and movement amplitudes of people will all interfere with face recognition and action analysis.
[0003] Currently, to improve the accuracy of face recognition, a sliding window is usually used to scan for preliminary localization of the face region, and then facial features are extracted through key point detection. However, in such methods, the sliding window needs to traverse a large number of invalid regions, and in scenarios such as vehicle fatigue warning, the requirement for recognition efficiency is relatively high, and the method of a large amount of redundant analysis is difficult to meet the requirements of this scenario. In addition, existing systems usually perform similarity matching through global feature embedding and a preset label library, or directly output attribute labels through a multi-task network. However, in this method, it is generally based on static labels for recognition, and static labels cannot capture instantaneous actions, resulting in insufficient emotion recognition accuracy. And if action labels are added, there is a problem that the action labels and static attribute labels are processed independently, lacking cross-modal association. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a face recognition method and system based on a face multi-label algorithm.
[0005] The present invention adopts the following technical solutions:
[0006] On the one hand, the present invention provides a face recognition method based on a face multi-label algorithm, and the method includes:
[0007] Performing a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and constructing a global coverage path of face features based on the face key points;
[0008] Performing a global scan on the current face region according to the global coverage path of face features, and matching with a preset label image according to the scan result to obtain an image global label corresponding to the current face image; wherein, the image global label includes: an action label and a static attribute label;
[0009] Determining the specified face key points corresponding to the action label, and determining the local ROI regions of each specified face key point in the current face image according to the neighborhoods corresponding to each specified face key point;
[0010] Perform multi-layer feature resampling on the local ROI region based on multiple frames of images corresponding to the current face image, so as to correct the action label according to the local action label corresponding to the sampling result;
[0011] Match the corrected action label, the static attribute label and the preset label result knowledge graph to implement face recognition of the current face image.
[0012] Optionally, in one or more embodiments of the present specification, perform a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and construct a global coverage path of face features based on the face key points, specifically including:
[0013] Determine the size range of the sliding window based on the image data and recognition scenario of the current face image, and perform a preliminary scan of the current face image according to the sliding windows of each size range to obtain face key points;
[0014] Classify each of the face key points based on the physiological structure features and functional features corresponding to each of the face key points to obtain multiple sets of face key point sets;
[0015] Determine the set path order corresponding to the face key point set according to the matching relationship between each face key point set and the preset physiological cognition order;
[0016] Determine the local path between each of the face key points within each face key point set based on the spatial position relationship corresponding to each of the face key points within each face key point set;
[0017] Determine whether the face key point set corresponding to the local path corresponds to a high-dynamic structure. If so, expand the branches of the local path based on the preset value corresponding to the high-dynamic structure to obtain an expanded local path;
[0018] Determine an initial global coverage path of face features according to the set path order, the local path between each of the face key points within the face key point set, and the expanded local path;
[0019] Perform smoothing processing on the initial global coverage path of face features to obtain a global coverage path of face features.
[0020] Optionally, in one or more embodiments of the present specification, the determining the size range of the sliding window based on the image data and recognition scenario of the current face image, and performing a preliminary scan of the current face image according to the sliding windows of each size range to obtain face key points, specifically includes:
[0021] Obtain the image attributes of the current face image to determine the image attribute tags corresponding to the current face image based on the image data, and obtain historical face images corresponding to the recognition scenario of the current face image and the image attribute tags; wherein, the image attributes include: light intensity, average face occupancy, background complexity;
[0022] Determine the initial size range of the sliding window corresponding to the current face image according to the sliding window size corresponding to the historical face image, and expand the initial size range based on the image resolution corresponding to the current face image to obtain the size range of the sliding window;
[0023] Based on the image size and the image resolution of the current face image, determine the corresponding number of image layers, and determine the corresponding scaling factor based on the preset recognition requirement information, so as to determine the number of sampleable layers of the current face image according to the number of image layers and the scaling factor;
[0024] Based on the scaling factor and the size range of the sliding window, determine the current size range of the sliding window for each sampleable layer of the current face image;
[0025] Perform sliding detection on the current face image of each sampleable layer based on the sliding window of the current size range to obtain face key points.
[0026] Optionally, in one or more embodiments of this specification, perform a global scan on the current face area according to the global coverage path of the face features, and match the scan result with a preset label image to obtain the image global label corresponding to the current face image, specifically including:
[0027] Perform a global scan on the current face area based on the global coverage path of the face features to obtain the basic data of each key point in the current face area and the key point displacement data between adjacent frames; wherein, the basic data includes: pixel value, position data;
[0028] Determine the static attribute features of the current face area according to the basic data of each key point in the current face area, and determine the action features of the current face area according to the key point displacement data between adjacent frames;
[0029] Determine the static attribute label corresponding to the static attribute features by performing similarity matching between the static attribute features and the face geometric features corresponding to each label in the preset static label image;
[0030] Determine the action label corresponding to the action features by performing similarity matching between the action features and the action trajectory features corresponding to each label in the preset action label image.
[0031] Optionally, in one or more embodiments of this specification, determine the specified facial key points corresponding to the action label, and determine the local ROI regions of each specified facial key point in the current facial image according to the neighborhoods corresponding to each specified facial key point. Specifically, it includes:
[0032] Based on the action features corresponding to the action label, determine the facial key points corresponding to the action label as the specified facial key points;
[0033] Taking each of the specified facial key points as the center, obtain the neighborhoods corresponding to each of the specified facial key points;
[0034] Determine whether the neighborhoods corresponding to each of the specified facial key points overlap;
[0035] If there is an overlap and the overlap ratio is greater than the preset ratio, then merge the neighborhoods to obtain the processed neighborhood as the local ROI region of the current facial image.
[0036] Optionally, in one or more embodiments of this specification, based on multiple frames of images corresponding to the current facial image, perform multi-layer feature resampling on the local ROI region. Specifically, it includes:
[0037] Based on the positions of the key points in the multiple frames of images corresponding to the current facial image and the positions of the key points in the local ROI region, align the spatial positions of the multiple frames of images and the local ROI region in the time series;
[0038] Determine the dynamic local ROI regions corresponding to the key points of the local ROI region in the aligned multiple frames of images;
[0039] Input the dynamic local ROI region into the preset CNN model to obtain multi-layer features output by different network levels of the preset CNN model;
[0040] Unify the spatial resolutions corresponding to the multi-layer features to fuse the unified multi-layer features to obtain the sampling result.
[0041] Optionally, in one or more embodiments of this specification, correct the action label according to the local action label corresponding to the sampling result. Specifically, it includes:
[0042] Input the sampling result into the classifier of the preset CNN model to obtain the local action label corresponding to the sampling result;
[0043] Compare the local action label of the multiple frames of images with the action label to determine whether there is a conflict between the local action label and the action label;
[0044] If so, based on the local action label, the action label is overwritten and modified to obtain a corrected action label.
[0045] Optionally, in one or more embodiments of this specification, before matching the corrected action label with the static attribute label and the pre-set label result knowledge graph to implement face recognition of the current face image, the method further includes:
[0046] Based on the existing face database and the existing behavior database, structured annotation data is obtained, and based on multiple public texts, action description and action attribute association rules are extracted to obtain unstructured data;
[0047] Entity extraction is performed on the unstructured data to obtain key entities and attribute information; wherein, the key entities include: action subject, action object;
[0048] Based on the relationship between the key entities, an initial label result knowledge graph is constructed, and the attribute information is added to the initial label result knowledge graph;
[0049] The entities corresponding to the structured annotation data are aligned with the initial label result knowledge graph, so as to expand the initial label result knowledge graph based on the structured annotation data to obtain a pre-set label result knowledge graph.
[0050] Optionally, in one or more embodiments of this specification, matching the corrected action label with the static attribute label and the pre-set label result knowledge graph to implement face recognition of the current face image specifically includes:
[0051] Determine the encoding format of the pre-set label result knowledge graph, so as to perform standardization processing on the corrected action label and the static attribute label based on the encoding format to obtain a processed action label and a processed static attribute label;
[0052] In the pre-set label result knowledge graph, search for entities and relationships corresponding to the processed action label and the processed static attribute label to obtain multiple matching results;
[0053] Obtain the similarity corresponding to each of the matching results, so as to sort each of the matching results based on the similarity, and use the matching result with the highest score as the face recognition result of the current face image.
[0054] On the other hand, the present invention also provides a face recognition system based on a face multi-label algorithm, and the system includes:
[0055] A construction unit for performing a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and constructing a global coverage path of face features based on the face key points;
[0056] An acquisition unit for globally scanning the current face area according to the global coverage path of the face features, so as to match the scanning result with a preset label image to obtain an image global label corresponding to the current face image; wherein, the image global label includes: an action label and a static attribute label;
[0057] A determination unit for determining the specified face key points corresponding to the action label, and determining the local ROI areas of each specified face key point in the current face image according to the grid areas corresponding to each specified face key point;
[0058] A correction unit for performing multi-layer feature resampling on the local ROI area based on multiple frames of images corresponding to the current face image, and correcting the action label according to the local action label corresponding to the sampling result;
[0059] An identification unit for matching the corrected action label and the static attribute label with a preset label result knowledge graph to implement face recognition of the current face image.
[0060] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0061] By performing a preliminary scan through a sliding window to obtain face key points and constructing a global coverage path, the overall and detailed features of the face can be comprehensively captured, avoiding missing important information and improving the accuracy of feature extraction. By performing a global scan based on the global coverage path and matching it with a preset label image, an image global label including an action label and a static attribute label can be obtained, accurately describing the features and attributes of the face image. Determining the specified face key points corresponding to the action label and the local ROI areas of their neighborhoods can focus on the key parts, reduce the interference of irrelevant information, and improve the pertinence and efficiency of recognition. Performing multi-layer feature resampling on the local ROI area and correcting the action label accordingly can further refine and optimize the description of face actions and improve the accuracy of action recognition. Matching the corrected action label and static attribute label with a preset label result knowledge graph and realizing face recognition by integrating multi-faceted information improve the accuracy and reliability of recognition and also enhance the adaptability to complex scenarios. Description of the Drawings
[0062] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. In the accompanying drawings:
[0063] Figure 1 It is a method flow chart of a face recognition method based on a face multi-label algorithm provided by the present invention;
[0064] Figure 2 It is a schematic structural diagram of a face recognition system based on a face multi-label algorithm provided by the present invention. Detailed implementation manners
[0065] The embodiments of this specification provide a face recognition method based on a face multi-label algorithm.
[0066] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0067] The following will detail the method in the present invention through the accompanying drawings.
[0068] Figure 1 It is a schematic diagram of the method flow of a face recognition method based on a face multi-label algorithm provided by the present invention. As Figure 1 shown, the method in the present invention at least includes the following execution steps:
[0069] S101: Perform a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and construct a global coverage path of face features based on the face key points.
[0070] In an actual scenario, such as in the scenario of fatigue driving recognition, the driving environment is complex and variable. The in-vehicle instrument panel, seats, etc. constitute a complex background, and the face images captured by the camera are often interfered by it. At this time, when the traditional method uses a sliding window to scan for preliminary face region localization and then extracts facial features through key point detection, since the sliding window needs to traverse a large number of invalid regions, especially in complex backgrounds or multi-person scenarios, the efficiency is low, and it is difficult to accurately locate the face in such a complex environment, resulting in deviation in feature extraction. Therefore, to solve this problem, in the embodiments of this specification, a preliminary scan is performed based on the sliding window corresponding to the current face image to obtain face key points, such as obtaining face key points like eyes, nose, mouth, etc. And a global coverage path of face features is constructed based on the obtained face key points, which can connect the scattered key points to form a complete path that can comprehensively describe the face features. This path not only contains the position information of each key point, but also reflects the spatial relationship and relative position between them. In the process of face recognition in the fatigue driving scenario, it can help to comprehensively capture the overall features and detailed features of the face. For example, by judging the position change of the eye key points, it can be determined whether there are fatigue signs such as drooping eyelids and frequent blinking, and whether there is a yawning action can be recognized based on the state of the mouth key point, greatly improving the accuracy and stability of fatigue driving recognition.
[0071] Specifically, in one or more embodiments of this specification, a preliminary scan is performed based on the sliding window corresponding to the current face image to obtain face key points, and a global coverage path of face features is constructed based on the face key points, which specifically includes the following processes:
[0072] First, based on the image data of the current face image and the recognition scenario, determine the size range of the sliding window, and perform a preliminary scan of the current face image according to the sliding windows of each size range to obtain face key points. That is, determine the size range of the sliding window according to the image data of the current face image and the recognition scenario. For example, in an image with strong light, simple background and large face proportion, a larger-size sliding window can be selected, which can quickly scan the face region; while in an image with dim light, complex background and small face proportion, a smaller-size sliding window needs to be used to search for the face more carefully to ensure that key information is not missed. After determining the size range, a preliminary scan of the current face image can be performed based on the sliding windows of each size range to obtain face key points. According to the physiological structure features corresponding to the face key points, such as being located in different parts like eyes, nose, mouth, etc., and functional features, such as the key points related to eyes are used to express the look in the eyes, and those related to the mouth are used for speaking expressions, etc., the obtained face key points are classified into multiple groups of face key point sets. For example, all the key points related to eyes are grouped into one group, and those related to the mouth are grouped into another group.
[0073] Then, according to the matching relationship between each of the face key point sets and the preset physiological recognition order, the set path order corresponding to the face key point set is determined. It should be noted that the preset physiological recognition order refers to the scanning priority rule predefined according to the anatomical structure of the face, the functional importance and the visual recognition habits of human beings. Then, for each set of face key point sets, the local path between them is determined according to the spatial position relationship of each key point in the set. For example, in the set of key points of the eyes, the connection order and mode between key points such as the corners of the eyes and the eyeballs are determined. Such a local path reflects the relative position relationship of the key points in the set. Then, it is determined whether the face key point set corresponding to the local path corresponds to a high dynamic structure. If so, the branch of the local path is expanded based on the preset value corresponding to the high dynamic structure to obtain the expanded local path. It can be understood that the high dynamic structure corresponds to a part with a large amplitude or frequent movement, such as a mouth with a large shape change under different expressions such as talking and smiling. By expanding the local path branch, its dynamic changes can be described more accurately, so that the path can more accurately reflect the characteristic changes of these parts. Then, the initial facial feature global coverage path is determined according to the set path sequence, the local path between each facial key point in the facial key point set, and the expanded local path. The initial facial feature global coverage path is smoothed to obtain the facial feature global coverage path.
[0074] In the above process, the sliding window size range is determined according to the image data and the recognition scene, which can effectively cope with various complex image environments. Regardless of whether factors such as lighting, background, or face ratio change, the key points of the face can be accurately obtained by adjusting the window size, ensuring the versatility and stability of the method. The key points of the face are classified, and the path is constructed in combination with the physiological cognitive order and spatial position relationship. At the same time, the expansion of the high dynamic structure is considered to comprehensively and meticulously describe the characteristics of the face. Not only does it cover the static structural features of the face, but it also fully considers the dynamic change characteristics, so that the constructed path can more realistically reflect the full picture of the face.
[0075] Further, in one or more embodiments of the present specification, based on the image data of the current face image and the recognition scene, the size range of the sliding window is determined, so as to perform a preliminary scan on the current face image according to the sliding windows of each size range to obtain the face key points, specifically including:
[0076] First, obtain the image attributes of the current face image, so as to determine the image attribute tags corresponding to the current face image according to the image data. At the same time, obtain the historical face images corresponding to the recognition scenario and image attribute tags of the current face image. It should be noted that the image attributes include: light intensity, average face proportion, and background complexity. Then, according to the sliding window size corresponding to the historical face image, determine the initial size range of the sliding window corresponding to the current face image, and expand the initial size range based on the image resolution corresponding to the current face image to obtain the size range of the sliding window. That is to say, since different image resolutions will affect the number of pixels and the area size actually covered by the window, it is necessary to expand the initial size range in combination with the resolution of the current face image to obtain a sliding window size range more suitable for the current image. For example, an image with a high resolution may require a larger window size to cover sufficient feature regions, while an image with a low resolution requires a relatively smaller window to avoid over-sampling. Then, based on the image size and image resolution of the current face image, determine the corresponding number of image layers, and determine the corresponding scaling factor based on the preset recognition requirement information. Then, determine the number of sampleable layers of the current face image according to the number of image layers and the scaling factor. Based on the scaling factor and the size range of the sliding window, determine the current size range of the sliding window for each sampleable layer of the current face image. That is, because the image resolutions of different sampling layers are different, the size of the sliding window also needs to be adjusted accordingly to ensure that the image can be effectively scanned and key information can be obtained at each sampling layer. Use the sliding window with the determined size range to perform sliding detection on the current face image at each sampleable layer to obtain face key points.
[0077] In the above process, by combining image attributes, recognition scenarios, and historical images to determine the size range of the sliding window, the characteristics of the current image and past experience can be fully considered, making the sliding window size more adaptable to the current face image. Using sliding windows of appropriate sizes for detection at sampleable layers with different resolutions can take into account both the overall features and local details of the image, obtain face key points more comprehensively and accurately, thereby improving the accuracy of face key point detection and providing a more reliable data basis for subsequent tasks such as face recognition. In addition, determining the scaling factor and the number of sampleable layers according to the preset recognition requirement information can reasonably allocate computing resources while meeting the recognition accuracy requirements.
[0078] S102: Perform a global scan of the current face area according to the global coverage path of the face features, so as to match the scan result with the preset label image to obtain the global image label corresponding to the current face image; wherein, the global image label includes: action label and static attribute label.
[0079] In traditional face recognition technology, similarity matching is performed through global feature embedding and a pre-set label library, or attribute labels are directly output through a multi-task network. However, in this method, the recognition is generally based on static labels, and static labels cannot capture instantaneous actions, resulting in insufficient emotion recognition accuracy. Therefore, in the embodiments of this specification, in order to achieve a more comprehensive description of the face and make up for the deficiencies of traditional technologies in the information acquisition dimension. In the embodiments of this specification, the current face area will be globally scanned according to the global coverage path of face features, and then matched with the pre-set label image according to the scanning result to obtain the global image label corresponding to the current face image. Among them, it should be noted that: the global image label includes: action labels and static attribute labels. Through the method of global scanning and matching with the pre-set label image, the action labels and static attribute labels of the face can be obtained simultaneously, realizing a more comprehensive description of the face and making up for the deficiencies of traditional technologies in the information acquisition dimension. Using the global coverage path of face features for global scanning can capture the features of the face in various states more comprehensively. Combined with the pre-set label image matching, it can effectively handle the face information recognition in complex scenarios, improving the adaptability and accuracy of the face recognition system in complex environments.
[0080] Specifically, in one or more embodiments of this specification, the current face area is globally scanned according to the global coverage path of face features to match the scanning result with the pre-set label image to obtain the global image label corresponding to the current face image. The specific process includes the following:
[0081] The constructed global coverage path of facial features is used to comprehensively scan the current facial region, so as to obtain the basic data of each key point in the current facial region according to the scanning results, as well as the displacement data of key points between adjacent frames. It should be noted that the basic data includes pixel values and position data. That is to say, during the scanning process, two types of important data of each key point in the current facial region are obtained. One type is the basic data, including pixel values and position data. The pixel value reflects the image color information at the position where the key point is located, and the position data accurately records the coordinates of the key point in the image. These basic data describe the characteristics of facial key points from a static perspective. The other type is the displacement data of key points between adjacent frames, which records the position change of the facial key points in the current frame relative to the previous frame in the video stream and is used to reflect the dynamic characteristics of the face. After obtaining the basic data of each key point in the current facial region and the displacement data of key points between adjacent frames, the static attribute characteristics of the current facial region, such as gender, approximate age range, skin color, etc., will be determined according to the basic data of each key point in the current facial region, and the action characteristics of the current facial region, such as smiling, nodding, turning the head, etc., will be determined according to the displacement data of key points between adjacent frames. The extracted static attribute characteristics are matched with the facial geometric characteristics corresponding to each label in the preset static label image to find the preset label with the highest similarity, so as to determine the static attribute label corresponding to the current facial static attribute characteristics. The determined action characteristics are matched with the action characteristics corresponding to each label in the preset action label image to find the most matching preset action label, and then the action label corresponding to the current facial action characteristics is determined.
[0082] In the above steps, not only the static characteristics of the face are concerned, but also the dynamic action characteristics are obtained through the displacement data of key points between adjacent frames, comprehensively covering the static attributes and action information of the face, providing rich data support for subsequent face recognition and analysis. Compared with the traditional recognition method based only on static facial features, it can describe the face state more comprehensively. By combining static attribute characteristics and action characteristics for matching, the face information can be confirmed from multiple angles. Even if some static characteristics are partially missing due to occlusion or other reasons, the action characteristics can still be used to assist in recognition, thereby improving the accuracy and stability of face recognition in complex scenarios. The static attribute characteristics and action characteristics are separately matched. This feature-based matching method is more refined and accurate. Different characteristics have their corresponding matching strategies and label libraries, which can be optimized for different types of information, avoiding interference between different types of characteristics and improving the reliability and credibility of the matching.
[0083] S103: Determine the specified facial key points corresponding to the action label, and determine the local ROI region of each specified facial key point in the current facial image according to the neighborhood corresponding to each specified facial key point.
[0084] Traditional face recognition and action analysis often take the entire face as the analysis object, ignoring the characteristics that different actions correspond to specific key regions, resulting in low analysis accuracy. Therefore, in order to improve the recognition accuracy, in the embodiments of this specification, the designated face key points corresponding to the action labels will be determined, and then the local ROI regions of each designated face key point in the current face image will be determined according to the neighborhoods corresponding to each designated face key point.
[0085] Specifically, in one or more embodiments of this specification, to determine the designated face key points corresponding to the action labels and determine the local ROI regions of each designated face key point in the current face image according to the neighborhoods corresponding to each designated face key point, it specifically includes:
[0086] Different action labels correspond to specific face action features. Therefore, based on the action features corresponding to the action labels, the face key points corresponding to the action labels are determined as the designated face key points. Then, with each designated face key point as the center, the neighborhoods corresponding to each designated face key point are obtained. Among them, the neighborhood contains the surrounding pixel information around the key point, and its size and shape can be set according to actual needs. For example, it can be set as a circular area with the key point as the center and a certain radius, or a square area with a fixed side length. The purpose of obtaining the neighborhood is to collect the surrounding information closely related to the key point, which helps to more comprehensively understand the changes of the key point and the performance of the action in the local area. Then, after determining the neighborhoods corresponding to each designated face key point, check whether there is an overlapping part between these neighborhoods and calculate the proportion of the overlapping area. When it is found that the neighborhoods overlap and the overlapping proportion is greater than the preset proportion, these neighborhoods are merged to obtain the processed neighborhood as the local ROI region of the current face image.
[0087] In the above process, by determining the designated face key points and their neighborhoods based on the action features, the regions closely related to specific actions can be accurately found, avoiding non-targeted processing of the entire face image and making the analysis more focused on the key parts where the action occurs. When determining the local ROI region, merging the overlapping neighborhoods avoids repeated processing of redundant information. This not only reduces the data processing volume, lowers the computational complexity, but also avoids the error accumulation that may be caused by repeated processing of the overlapping regions, improving the operating efficiency of the system and the reliability of the analysis results.
[0088] S104: Based on multiple frames of images corresponding to the current face image, perform multi-layer feature resampling on the local ROI region to correct the action label according to the local action label corresponding to the sampling result.
[0089] The action information obtained from a single-frame image is limited and it is difficult to comprehensively and accurately reflect the real action. Therefore, to improve the recognition accuracy, multi-layer feature resampling is performed on the local ROI region according to multiple frames of images corresponding to the current face image, so as to correct the action label according to the local action label corresponding to the sampling result. Specifically, in one or more embodiments of this specification, multi-layer feature resampling is performed on the local ROI region based on multiple frames of images corresponding to the current face image, which specifically includes:
[0090] Since the face may have slight movement and rotation in a continuous video frame, to ensure the consistency of multiple frames of images and the local ROI region in time and space, so that subsequent analysis can be carried out based on a unified coordinate system. The spatial positions of multiple frames of images and the local ROI region are aligned in the dimension of time series by comparing the positions of key points in multiple frames of images with the positions of key points in the local ROI region. After completing the spatial position alignment, according to the part of the multiple frames of images corresponding to the key points of the local ROI region, the corresponding dynamic local ROI region in each frame of image is determined. Since the face pose, expression, etc. in multiple frames of images may change, the local ROI region will also change dynamically accordingly. Therefore, it is necessary to accurately find the region corresponding to the key points of the initial local ROI region in each frame of image to form a dynamic local ROI region. These dynamic local ROI regions can reflect the local feature changes of the face at different times. The determined dynamic local ROI regions are input into a pre-set convolutional neural network model. At this time, different network levels of the pre-set CNN model can extract features of different levels and abstraction degrees, and multi-layer features output by different network levels of the pre-set CNN model are obtained. Since the multi-layer features output by different network levels may have different spatial resolutions, to effectively fuse these features, it is necessary to unify their spatial resolutions, and then fuse the unified multi-layer features to obtain the sampling result.
[0091] In this process, through spatial position alignment, the consistency of multiple frames of images and the local ROI region in time and space is ensured, so that the features extracted subsequently can accurately reflect the local changes of the face at different times. The setting of the dynamic local ROI region enables the key local regions to be concerned at different times, improving the ability to capture dynamic changes. In addition, by unifying the spatial resolutions and performing fusion, these different-level feature information can be integrated to form a more representative sampling result. This comprehensive feature can better reflect the overall features and changes of the face locally, contributing to improving the performance of subsequent recognition. And by processing multiple frames of images and fusing multi-layer features, features at different levels can complement each other, making subsequent face recognition more accurate.
[0092] Further, in one or more embodiments of this specification, the action label is corrected according to the local action label corresponding to the sampling result, specifically including:
[0093] The sampling result obtained after multi-layer feature resampling is input into the classifier of the pre-set CNN model, and the local action label corresponding to the sampling result can be obtained based on the analysis and matching of the classifier. Since the preliminarily determined action label is obtained based on the analysis of some basic features in the previous processing flow, it may be inaccurate or incomplete. Therefore, the local action labels of multiple frames of images are compared with the action label to determine whether there is a conflict between the local action label and the action label. For example, if the preliminary action label is judged as "normal expression" while the local action label shows "slight frown", this indicates a conflict between the two. If a conflict is found between the local action label and the action label during the comparison process, the preliminary action label is overwritten and modified with the local action label. It can be understood that the local action label is obtained based on the more refined multi-layer feature resampling and classifier judgment of multiple frames of images, and contains richer and more accurate information. Through the overwrite modification, the corrected action label is obtained, making it more consistent with the actual action situation of the face.
[0094] S105: Match the corrected action label, the static attribute label, and the pre-set label result knowledge graph to implement the face recognition of the current face image.
[0095] Traditional face recognition mainly relies on the static features of the face, such as the shape and position of facial features. However, the face has rich dynamic changes in different scenarios, which will cause great changes in the static features, thereby reducing the accuracy of recognition. Therefore, in this application, by combining the corrected action label and the static attribute label for matching to implement the face recognition of the current face image, the face features can be described from multiple dimensions, increasing the basis for recognition. It makes up for the deficiency of single static feature recognition and improves the accuracy of recognition.
[0096] Further, in one or more embodiments of this specification, before matching the corrected action label, the static attribute label, and the pre-set label result knowledge graph to implement the face recognition of the current face image, the method further includes the following process:
[0097] First, obtain structured annotation data from the existing face database and the existing behavior database, and extract the association rules between action descriptions and action attributes based on multiple public texts to obtain unstructured data. It can be understood that the public texts may contain various descriptions of actions. By analyzing these texts, the connections between actions and related attributes can be extracted. For example, the action of "smiling" may be associated with the attribute of "friendly". Then, perform entity extraction on the unstructured data to identify key entities and attribute information. Among them, the key entities include: the action subject and the action object. Then, construct an initial label result knowledge graph based on the key entities and the relationships between them, and add the extracted attribute information to the initial label result knowledge graph to make the knowledge graph more rich and specific. Align the entities corresponding to the structured annotation data with the initial label result knowledge graph. In this way, use the rich information in the structured annotation data to expand the initial knowledge graph and obtain a preset label result knowledge graph.
[0098] This process integrates structured and unstructured data and makes full use of the data advantages of different sources. Structured annotation data provides accurate and standardized information, while unstructured data contains richer, more diverse text descriptions and potential association rules. The combination of the two enables the knowledge graph to cover more comprehensive information. Through entity extraction and relationship construction, key knowledge can be effectively extracted from complex data and presented intuitively and clearly in the form of a knowledge graph. In addition, the method of first constructing an initial knowledge graph and then expanding it by aligning with structured annotation data makes the knowledge graph highly scalable and flexible.
[0099] Specifically, in one or more embodiments of this specification, match the corrected action label and the static attribute label with the preset label result knowledge graph to implement face recognition of the current face image, specifically including:
[0100] Determine the encoding format of the preset label result knowledge graph to standardize the corrected action label and the static attribute label based on the encoding format, and obtain the processed action label and the processed static attribute label. After the standardization process, search for the entities and relationships corresponding to the processed action label and the static attribute label in the preset label result knowledge graph to obtain multiple matching results. For the multiple matching results found, calculate the similarity corresponding to each result, and sort the matching results based on the similarity. Take the matching result with the highest score as the face recognition result of the current face image.
[0101] By determining the encoding format and performing standardization processing, the compatibility and consistency of the corrected action tags and static attribute tags with the data of the preset tag result knowledge graph are ensured. This helps to improve the accuracy and efficiency of data matching, avoid errors or omissions caused by inconsistent data formats, and enables the effective integration and utilization of data from different sources under the framework of the knowledge graph. Calculating the similarity of the matching results and sorting them can provide a quantitative basis for the recognition results, avoid the uncertainty of subjective judgment, and further improve the accuracy of recognition.
[0102] Based on the same inventive concept, the present invention also provides a face recognition system based on a face multi-label algorithm, and its structure is as Figure 2 shown.
[0103] Figure 2 It is a schematic structural diagram of a face recognition system based on a face multi-label algorithm provided by the present invention. As Figure 2 shown, the system in the present invention specifically includes:
[0104] A construction unit 201, configured to perform a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and construct a global face feature coverage path based on the face key points;
[0105] An acquisition unit 202, configured to perform a global scan on the current face area according to the global face feature coverage path, and match the scan result with a preset label image to obtain an image global label corresponding to the current face image; wherein, the image global label includes: an action label and a static attribute label;
[0106] A determination unit 203, configured to determine the specified face key points corresponding to the action label, and determine the local ROI areas of each specified face key point in the current face image according to the grid areas corresponding to the specified face key points;
[0107] A correction unit 204, configured to perform multi-layer feature resampling on the local ROI area based on multiple frames of images corresponding to the current face image, and correct the action label according to the local action label corresponding to the sampling result;
[0108] An identification unit 205, configured to match the corrected action label and the static attribute label with the preset label result knowledge graph to implement face recognition of the current face image.
[0109] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0110] The device provided by the present invention corresponds to the method one by one. Therefore, the device also has beneficial technical effects similar to those of its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device will not be elaborated here.
[0111] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0112] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0113] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A face recognition method based on a face multi-label algorithm, characterized in that, The method includes: Performing a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and constructing a global coverage path of face features based on the face key points; Performing a global scan on the current face area according to the global coverage path of face features, so as to match with a preset label image according to the scan result to obtain an image global label corresponding to the current face image; wherein, the image global label includes: an action label and a static attribute label; Determining specified face key points corresponding to the action label, so as to determine local ROI regions of each specified face key point in the current face image according to the neighborhoods corresponding to the specified face key points; Performing multi-layer feature resampling on the local ROI regions based on multiple frames of images corresponding to the current face image, so as to correct the action label according to the local action label corresponding to the sampling result; Matching the corrected action label, the static attribute label and a preset label result knowledge graph to implement face recognition of the current face image.
2. The face recognition method based on a multi-label face algorithm according to claim 1, characterized in that, Performing a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and constructing a global coverage path of face features based on the face key points, specifically including: Determining the size range of the sliding window based on the image data and recognition scenario of the current face image, so as to perform a preliminary scan on the current face image according to the sliding windows of each size range to obtain face key points; Classifying each of the face key points based on the physiological structure features and functional features corresponding to the face key points to obtain multiple sets of face key point sets; Determining the set path order corresponding to the face key point set according to the matching relationship between each face key point set and a preset physiological recognition order; Determining local paths between the face key points within each face key point set based on the spatial position relationships corresponding to the face key points within each face key point set; Determining whether the face key point set corresponding to the local path corresponds to a high-dynamic structure, and if so, expanding the branches of the local path based on a preset value corresponding to the high-dynamic structure to obtain an expanded local path; Determining an initial global coverage path of face features according to the set path order, the local paths between the face key points within the face key point set, and the expanded local path; Performing smoothing processing on the initial global coverage path of face features to obtain a global coverage path of face features.
3. The face recognition method based on a multi-label face algorithm according to claim 2, characterized in that, The determining the size range of the sliding window based on the image data and recognition scenario of the current face image, so as to perform a preliminary scan on the current face image according to the sliding windows of each size range to obtain face key points, specifically including: Obtaining the image attributes of the current face image to determine an image attribute label corresponding to the current face image based on the image data, and obtaining a historical face image corresponding to the recognition scenario of the current face image and the image attribute label; wherein, the image attributes include: light intensity, average face ratio, background complexity; Determine the initial size range of the sliding window corresponding to the current face image according to the size of the sliding window corresponding to the historical face image, and expand the initial size range based on the image resolution corresponding to the current face image to obtain the size range of the sliding window; Determine the corresponding number of image layers based on the image size and the image resolution of the current face image, and determine the corresponding scaling factor based on the preset recognition requirement information, so as to determine the number of sampleable layers of the current face image according to the number of image layers and the scaling factor; Determine the current size range of the sliding window of each sampleable layer of the current face image based on the scaling factor and the size range of the sliding window; Perform sliding detection on the current face image of each sampleable layer based on the sliding window within the current size range to obtain face key points.
4. A face recognition method based on a face multi-label algorithm according to claim 1, wherein Perform a global scan on the current face area according to the global coverage path of the face features, so as to match the scan result with the preset label image to obtain the global image label corresponding to the current face image, specifically including: Perform a global scan on the current face area based on the global coverage path of the face features, so as to obtain the basic data of each key point in the current face area and the key point displacement data between adjacent frames based on the scan result; wherein, the basic data includes: pixel value, position data; Determine the static attribute features of the current face area according to the basic data of each key point in the current face area, and determine the action features of the current face area according to the key point displacement data between adjacent frames; Determine the static attribute label corresponding to the static attribute features by performing similarity matching between the static attribute features and the face geometric features corresponding to each label in the preset static label image; Determine the action label corresponding to the action features by performing similarity matching between the action features and the action trajectory features corresponding to each label in the preset action label image.
5. A face recognition method based on a face multi-label algorithm according to claim 4, characterized in that, Determine the specified face key points corresponding to the action label, so as to determine the local ROI area of each specified face key point in the current face image according to the neighborhood corresponding to each specified face key point, specifically including: Determine the face key points corresponding to the action label as the specified face key points based on the action features corresponding to the action label; Take each of the specified face key points as the center and obtain the neighborhood corresponding to each of the specified face key points; Determine whether the neighborhoods corresponding to each of the specified face key points overlap; If there is an overlap and the overlap ratio is greater than the preset ratio, merge the neighborhoods to obtain the processed neighborhood as the local ROI area of the current face image.
6. A face recognition method based on a face multi-label algorithm according to claim 1, characterized in that, Perform multi-layer feature resampling on the local ROI area based on multiple frames of images corresponding to the current face image, specifically including: Align the spatial positions of the multiple frames of images and the local ROI area in the time series based on the positions of each key point in the multiple frames of images corresponding to the current face image and the positions of each key point in the local ROI area; Determine a dynamic local ROI region corresponding to the key points of the local ROI region in the aligned multi-frame image; Input the dynamic local ROI region into a pre-set CNN model to obtain multi-level features output by different network levels of the pre-set CNN model; Unify the spatial resolutions corresponding to the multi-level features to fuse the unified multi-level features to obtain a sampling result.
7. A face recognition method based on a face multi-label algorithm according to claim 6, characterized in that Correct the action label according to the local action label corresponding to the sampling result, specifically including: Input the sampling result into the classifier of the pre-set CNN model to obtain the local action label corresponding to the sampling result; Compare the local action label of the multi-frame image with the action label to determine whether there is a conflict between the local action label and the action label; If so, overwrite and modify the action label based on the local action label to obtain a corrected action label.
8. The face recognition method based on a multi-label face algorithm according to claim 1, wherein Before matching the corrected action label, the static attribute label and the pre-set label result knowledge graph to implement face recognition of the current face image, the method further includes: Based on the existing face database and the existing behavior database, obtain structured annotation data, and extract action description and action attribute association rules from multiple public texts to obtain unstructured data; Perform entity extraction on the unstructured data to obtain key entities and attribute information; wherein, the key entities include: action subject, action object; Based on the relationship between the key entities, construct an initial label result knowledge graph, and add the attribute information to the initial label result knowledge graph; Align the entities corresponding to the structured annotation data with the initial label result knowledge graph to expand the initial label result knowledge graph based on the structured annotation data to obtain a pre-set label result knowledge graph.
9. A face recognition method based on a face multi-label algorithm according to claim 1, characterized in that, Match the corrected action label, the static attribute label and the pre-set label result knowledge graph to implement face recognition of the current face image, specifically including: Determine the encoding format of the pre-set label result knowledge graph to perform standardization processing on the corrected action label and the static attribute label based on the encoding format to obtain a processed action label and a processed static attribute label; In the pre-set label result knowledge graph, search for entities and relationships corresponding to the processed action label and the processed static attribute label to obtain multiple matching results; Obtain the similarity corresponding to each of the matching results, sort each of the matching results based on the similarity, and use the matching result with the highest score as the face recognition result of the current face image.
10. The face recognition system based on a face multi-label algorithm according to claim 1, characterized in that, The system includes: A construction unit for performing a preliminary scan based on a sliding window corresponding to the current face image to obtain face key points, and constructing a global coverage path of face features based on the face key points; An acquisition unit for globally scanning the current face area according to the global coverage path of the face features, so as to match the scanning result with a preset label image to obtain an image global label corresponding to the current face image; wherein, the image global label includes: an action label and a static attribute label; A determination unit for determining specified face key points corresponding to the action label, so as to determine local ROI areas of the specified face key points in the current face image according to the grid areas corresponding to the specified face key points; A correction unit for performing multi-layer feature resampling on the local ROI area based on multiple frames of images corresponding to the current face image, so as to correct the action label according to the local action label corresponding to the sampling result; An identification unit for matching the corrected action label with the static attribute label and a preset label result knowledge graph to implement face recognition of the current face image.
Citation Information
Patent Citations
Face attribute recognition system and method based on multiple areas of face
CN110069994A
Facial expression recognition method and system based on regional grouping and internal association fusion
CN112990007A
Virtual makeup trying method based on face recognition
CN113643397A
Facial recognition model training method and facial recognition method
CN117593776A
Character action recognition analysis method and system based on infrared laser and deep learning
CN118747911A