Intelligent collection and distributed storage management method of student photos based on face recognition

Through dynamic focus scanning and adaptive region expansion combined with facial key points and phase spectrum change characteristics, accurate collection and efficient storage of student photo records are achieved, solving the problems of inaccurate face positioning and low storage efficiency in traditional methods, and improving the acquisition efficiency and data management security.

CN120296185BActive Publication Date: 2025-08-19HANGZHOU RONGBO EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510750475.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-19
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional student photo collection lacks an intelligent judgment mechanism, resulting in frequent replays, inaccurate positioning, inefficient storage methods, and difficult to support the rapid retrieval and comparison of large-scale facial images.

Method used

Accurate face positioning is achieved through dynamic focus scanning and adaptive region expansion, real-time evaluation is performed by combining facial key points and phase spectrum change characteristics, shooting trigger signals are generated, and dual index features are extracted through pyramid decomposition for distributed storage.

Benefits of technology

It improves the accuracy and efficiency of facial photos collection, reduces the number of reshoots, reduces storage costs, and improves data access efficiency and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296185B_ABST
    Figure CN120296185B_ABST
Patent Text Reader

Abstract

This invention provides a method for intelligently collecting and distributing student photos based on facial recognition, which relates to the field of facial recognition management technology. The method involves acquiring facial images, performing dynamic focus scanning to locate the face, extracting facial key points to calculate state assessment parameters, generating a capture trigger signal based on phase spectrum variation characteristics and stability prior information, collecting and enhancing a standard facial image, and extracting structural and texture fingerprints to form dual index features. This invention improves the intelligence of facial photo collection, enhances facial recognition accuracy, and achieves efficient distributed storage management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of face recognition management technology, and in particular to a method for intelligent collection and distributed storage management of student photos based on face recognition. Background Art

[0002] With the in-depth development of educational informatization, photo collection and management in student registration management systems have become an essential component of school management. Traditional student registration photo collection typically relies on manual photography, screening, and centralized storage, which is inefficient for large-scale student registration information management. In recent years, facial recognition technology has been widely applied in various fields, demonstrating its immense value in identity verification, security monitoring, and other areas. Applying facial recognition technology to the intelligent collection and management of student registration photos can significantly improve photo collection quality and management efficiency.

[0003] However, there are several problems in the collection and storage management of student photos. Traditional photo collection lacks an intelligent judgment mechanism and cannot evaluate the suitability of facial conditions in real time, resulting in photos often having problems such as closed eyes, tilted head, and inappropriate expressions, which require frequent reshoots, wasting time and resources; existing face positioning algorithms have low accuracy under complex backgrounds or lighting conditions, making it difficult to accurately extract standardized face areas, affecting subsequent processing and recognition effects; conventional photo storage methods often adopt a single centralized architecture, which not only has low retrieval efficiency, but is also prone to performance bottlenecks when the system load increases. At the same time, it lacks an effective feature indexing mechanism, making it difficult to support the rapid retrieval and comparison needs of large-scale facial images.

[0004] With the increase in the number of students and the improvement in the demand for student registration management, there is an urgent need for a technical solution that can intelligently collect high-quality photos and achieve efficient storage management to meet the needs of education management. Summary of the Invention

[0005] The embodiment of the present invention provides a method for intelligent collection and distributed storage management of student photos based on face recognition, which can solve the problems in the existing technology.

[0006] A first aspect of an embodiment of the present invention provides a method for intelligently collecting and distributing student photos based on face recognition, comprising:

[0007] Acquire a facial image video stream through a camera device, and preprocess it to obtain a preprocessed image sequence;

[0008] Performing dynamic focus scanning on the preprocessed image sequence, calculating the local texture complexity and neighborhood grayscale uniformity of the scan position to obtain regional interest, performing adaptive region expansion with the highest point of regional interest as the center to obtain face positioning parameters, and extracting the target face image;

[0009] Based on the target face image, facial key points are extracted according to preset points, and state evaluation parameters including facial posture angle value, eye opening value and expression parameter value are calculated based on the facial key points;

[0010] The phase spectrum change characteristics of continuous target face images are calculated, the state stability score is determined based on the stability prior information, and the shooting trigger signal is generated based on the adaptive threshold.

[0011] In response to the shooting trigger signal, the face photo is collected and enhanced to obtain a standard face image;

[0012] The multi-scale features of facial images are extracted based on the pyramid decomposition structure, and the structural fingerprint and texture fingerprint are generated through frequency domain transformation compression, which are then fused in the feature manifold space to form a dual index feature.

[0013] The standard face image and dual index features are stored in a distributed storage system to generate an electronic certificate containing the storage location.

[0014] In an optional embodiment, performing dynamic focus scanning on the pre-processed image sequence and calculating the local texture complexity and neighborhood grayscale uniformity at the scanning position to obtain the regional interest includes:

[0015] A dynamic angle step is established based on the image center coordinates of the preprocessed image sequence, a dynamic radius step is generated by multiplying the dynamic angle step by the dynamic expansion factor, and the initial scanning coordinates are calculated based on the dynamic angle step and the dynamic radius step;

[0016] Calculating a local response intensity value based on the initial scanning coordinates, generating a focus adjustment coefficient based on the ratio of the local response intensity value to a preset reference threshold, and adjusting the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position;

[0017] A first pixel window is selected based on the dynamic focus scanning position, and the local texture complexity is generated by calculating the cumulative difference between the pixels in the first pixel window and the weighted average value of the window. A second pixel window is selected based on the dynamic focus scanning position, and the neighborhood grayscale uniformity is generated by calculating the ratio of the pixel discreteness of the second pixel window to the overall discreteness of the image.

[0018] The local texture complexity and the neighborhood grayscale uniformity are multiplied by the corresponding preset weights to obtain the first weighted feature and the second weighted feature, respectively. The basic area feature value is determined, and the boundary attenuation coefficient is calculated according to the distance from the dynamic focus scanning position to the image boundary. The product of the basic area feature value and the boundary attenuation coefficient is used to generate the regional interest.

[0019] In an optional embodiment, adaptive region expansion is performed with the highest regional interest point as the center to obtain face location parameters, and extracting the target face image includes:

[0020] Determining a region growth starting point at a location with a maximum region interest, constructing a polar coordinate grid with the region growth starting point as the center, calculating an interest gradient value on each directional axis of the polar coordinate grid, multiplying the interest gradient value by the directional axis length to generate an axial expansion distance, determining boundary point coordinates based on the axial expansion distance, and performing a low-pass filter on the sequence of boundary point coordinates to obtain a smooth region outline;

[0021] Ellipse fitting is performed on the smooth area contour to calculate the area center coordinates, major axis length, minor axis length and direction angle, and combine them to generate a positioning parameter set. A coordinate transformation matrix is constructed according to the positioning parameter set. The coordinate transformation matrix is applied to the preprocessed image sequence for geometric mapping, and size normalization is performed to extract the target face image.

[0022] In an optional embodiment, calculating phase spectrum change features for continuous target facial images, determining a state stability score in combination with stability prior information, and generating a shooting trigger signal based on an adaptive threshold comprises:

[0023] Performing a two-dimensional Fourier transform on continuous target facial images to obtain a frequency domain representation, extracting a phase component from the frequency domain representation to generate a phase spectrum sequence, and stacking the phase spectrum sequence in the time dimension to form a phase spectrum tensor;

[0024] Calculating the phase spectrum differences of adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence, constructing a frequency weight function based on a Gaussian kernel, and performing a weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain a phase spectrum change feature;

[0025] Acquire stability prior information of the face area, map the stability prior information to the frequency space to obtain a state evaluation parameter, and generate a state stability score according to the combined distribution law of the phase spectrum change characteristics and the state evaluation parameter;

[0026] Calculate the mean of the state stability scores within the historical time window, determine the dynamic adjustment coefficient by the ratio of the mean to the preset stability threshold, and adjust the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive judgment threshold; when the state stability scores are all greater than the adaptive judgment threshold within a continuous preset frame number threshold, a shooting trigger signal is generated.

[0027] In an optional embodiment, obtaining the stability prior information of the face region includes:

[0028] Obtaining coordinates of facial feature points in a target facial image, calculating displacement vectors of the facial feature point coordinates between adjacent image frames, and constructing a facial posture motion model based on the displacement vectors;

[0029] Performing an affine transformation on the facial feature point coordinates to obtain reference plane facial coordinates, calculating a feature distance matrix of the reference plane facial coordinates, and extracting facial region deformation parameters;

[0030] Calculating the translation component, rotation component and scaling component in the face posture motion model respectively, and combining them with the face region deformation parameter to obtain a face motion feature vector;

[0031] Performing principal component decomposition on the facial motion feature vector to obtain a feature projection matrix, projecting the facial motion feature vector onto the feature projection matrix to obtain a reduced-dimensional feature representation, and constructing a stability evaluation space based on the reduced-dimensional feature representation;

[0032] A local optimal stability interval is extracted from the stability evaluation space, and statistical features of the local optimal stability interval are output as stability priori information of the face region.

[0033] In an optional embodiment, multi-scale features of a facial image are extracted based on a pyramid decomposition structure, structural fingerprints and texture fingerprints are generated through frequency domain transform compression, and dual index features are formed by fusing them in a feature manifold space. The following steps are included:

[0034] A pyramid decomposition structure is constructed for standard face images. The edge direction histogram is extracted at each scale level of the pyramid decomposition structure. The edge direction histograms of adjacent scale levels are cross-compared to obtain a multi-scale structural feature matrix.

[0035] Calculating grayscale co-occurrence features at each scale level of the pyramid decomposition structure, accumulating the grayscale co-occurrence features in different directions to form a directional accumulation curve, and generating a multi-scale texture feature matrix based on the distribution law of the directional accumulation curve;

[0036] Performing discrete cosine transform on the multi-scale structural feature matrix to obtain a frequency domain coefficient matrix, selecting low-frequency components in the frequency domain coefficient matrix to construct a structural feature descriptor, and generating a structural fingerprint code based on the energy distribution of the structural feature descriptor;

[0037] Performing wavelet transform on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, extracting high-frequency components in the wavelet coefficient matrix to construct a texture response sequence, and generating a texture fingerprint code based on the energy distribution of the texture response sequence;

[0038] The structural fingerprint code and the texture fingerprint code are mapped to the local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and a dual index feature is obtained by adaptive fusion based on feature alignment loss optimization.

[0039] In an optional embodiment, mapping the structural fingerprint and the texture fingerprint to a local feature manifold space, constructing a cross-modal correlation matrix in the local feature manifold space, and adaptively fusing the dual index features based on feature alignment loss optimization includes:

[0040] Calculate the feature distance distribution of the structural fingerprint code and texture fingerprint code in the local neighborhood respectively, obtain the structural similarity matrix and texture similarity matrix, and construct the feature distribution representation;

[0041] Performing cross-correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross-modal correlation matrix, calculating feature weight coefficients, and performing weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix;

[0042] Constructing a nonlinear mapping model including an encoding layer and a decoding layer, inputting a structural fingerprint code and a texture fingerprint code, calculating the alignment loss between the mapping features based on the manifold consistency matrix, and constructing an alignment loss function in combination with the reconstruction loss;

[0043] Based on the alignment loss function, loss iterative optimization is performed to obtain an optimized nonlinear mapping function, and the optimized nonlinear mapping function is used to remap the structural fingerprint code and the texture fingerprint code to obtain the remapped feature;

[0044] The similarity between the remapped features is calculated to obtain a feature similarity matrix, feature fusion weights are calculated in combination with the feature distribution representation, and the remapped features are weightedly combined according to the feature fusion weights to obtain dual index features.

[0045] A second aspect of an embodiment of the present invention provides an intelligent collection and distributed storage management system for student photos based on face recognition, including:

[0046] The first unit is used to obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence;

[0047] The second unit is configured to perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood grayscale uniformity of the scanning position to obtain regional interest, perform adaptive region expansion with the highest point of regional interest as the center to obtain face positioning parameters, and extract the target face image;

[0048] The third unit is used to extract facial key points according to preset points based on the target face image, and calculate state evaluation parameters including facial posture angle value, eye opening value and expression parameter value based on the facial key points;

[0049] The fourth unit is used to calculate the phase spectrum change characteristics of continuous target face images, determine the state stability score based on the stability prior information, and generate a shooting trigger signal based on the adaptive threshold;

[0050] The fifth unit is used to respond to the shooting trigger signal, collect the face photo and enhance the processing to obtain the standard face image;

[0051] The sixth unit is used to extract multi-scale features of facial images based on the pyramid decomposition structure, generate structural fingerprints and texture fingerprints through frequency domain transformation compression, and fuse them in the feature manifold space to form dual index features;

[0052] The seventh unit is used to store the standard face image and the dual index features into the distributed storage system and generate an electronic certificate including the storage location.

[0053] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0054] processor;

[0055] a memory for storing processor-executable instructions;

[0056] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0057] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0058] In an embodiment of the present invention, a method for intelligent collection and distributed storage management of student photos based on face recognition achieves precise face positioning through dynamic focus scanning and adaptive area expansion, effectively solving the problems of inaccurate positioning and susceptibility to environmental interference in traditional face collection, and improving the positioning accuracy and stability of face images; adopts facial key point extraction and state evaluation parameter calculation, combined with phase spectrum change characteristics and stability prior information, to achieve real-time evaluation of facial state and intelligent judgment of the best shooting time, ensuring that the collected face photos meet standard requirements, effectively reducing the number of manual intervention and repeated shooting, and improving collection efficiency; by extracting structural features and texture features and compressing them to form dual index features, combined with a distributed storage system, efficient retrieval and secure management of student photos are achieved, which not only reduces storage costs, but also improves data access efficiency and security, and provides reliable technical support for student information management. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a flow chart of a method for intelligent collection and distributed storage management of student photos based on face recognition according to an embodiment of the present invention;

[0060] Figure 2 Schematic diagram of the dynamic focus scanning and regional interest calculation process;

[0061] Figure 3 Schematic diagram of the performance comparison of face stability assessment methods. DETAILED DESCRIPTION

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0063] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0064] Figure 1 FIG is a flow chart of a method for intelligent collection and distributed storage management of student photos based on face recognition according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0065] Acquire a facial image video stream through a camera device, and preprocess it to obtain a preprocessed image sequence;

[0066] Performing dynamic focus scanning on the preprocessed image sequence, calculating the local texture complexity and neighborhood grayscale uniformity of the scan position to obtain regional interest, performing adaptive region expansion with the highest point of regional interest as the center to obtain face positioning parameters, and extracting the target face image;

[0067] Based on the target face image, facial key points are extracted according to preset points, and state evaluation parameters including facial posture angle value, eye opening value and expression parameter value are calculated based on the facial key points;

[0068] The phase spectrum change characteristics of continuous target face images are calculated, the state stability score is determined based on the stability prior information, and the shooting trigger signal is generated based on the adaptive threshold.

[0069] In response to the shooting trigger signal, the face photo is collected and enhanced to obtain a standard face image;

[0070] The multi-scale features of facial images are extracted based on the pyramid decomposition structure, and the structural fingerprint and texture fingerprint are generated through frequency domain transformation compression, which are then fused in the feature manifold space to form a dual index feature.

[0071] The standard face image and dual index features are stored in a distributed storage system to generate an electronic certificate containing the storage location.

[0072] In an optional embodiment, performing dynamic focus scanning on the pre-processed image sequence and calculating the local texture complexity and neighborhood grayscale uniformity of the scanning position to obtain the regional interest includes:

[0073] A dynamic angle step is established based on the image center coordinates of the preprocessed image sequence, a dynamic radius step is generated by multiplying the dynamic angle step by the dynamic expansion factor, and the initial scanning coordinates are calculated based on the dynamic angle step and the dynamic radius step;

[0074] Calculating a local response intensity value based on the initial scanning coordinates, generating a focus adjustment coefficient based on the ratio of the local response intensity value to a preset reference threshold, and adjusting the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position;

[0075] A first pixel window is selected based on the dynamic focus scanning position, and the local texture complexity is generated by calculating the cumulative difference between the pixels in the first pixel window and the weighted average value of the window. A second pixel window is selected based on the dynamic focus scanning position, and the neighborhood grayscale uniformity is generated by calculating the ratio of the pixel discreteness of the second pixel window to the overall discreteness of the image.

[0076] The local texture complexity and the neighborhood grayscale uniformity are multiplied by the corresponding preset weights to obtain the first weighted feature and the second weighted feature, respectively. The basic area feature value is determined, and the boundary attenuation coefficient is calculated according to the distance from the dynamic focus scanning position to the image boundary. The product of the basic area feature value and the boundary attenuation coefficient is used to generate the regional interest.

[0077] Figure 2 This is a flow chart of dynamic focus scanning and regional interest calculation, as shown in Figure 2As shown, in a specific embodiment, a dynamic angle step is established based on the coordinates of the center point of the image after the image sequence is preprocessed. Specifically, the preprocessing of the image sequence includes image denoising, contrast adjustment and color space conversion, converting the RGB image into a grayscale image, and removing noise by Gaussian filtering. The coordinates of the center point of the image can be set to half the width and height of the image. For example, for an image with a resolution of 640×480, the coordinates of the center point are (320, 240). The dynamic angle step is adaptively adjusted according to the complexity of the image. When the overall complexity of the image is high, the angle step is set to a smaller value, such as 5 degrees; when the overall complexity of the image is low, the angle step is set to a larger value, such as 15 degrees. The image complexity is quantified by calculating the grayscale co-occurrence matrix features of the image. The result value ranges from 0 to 1. The larger the value, the higher the image complexity.

[0078] The dynamic angle step is multiplied by the dynamic expansion factor to generate the dynamic radius step. The dynamic expansion factor is determined based on the brightness distribution of the image area. Areas with more uniform brightness have a larger expansion factor, such as 1.5; areas with drastic brightness changes have a smaller expansion factor, such as 0.8. In practical applications, the uniformity of the brightness distribution can be determined by calculating the standard deviation of the brightness of the local area of the image. When the standard deviation is less than 10, the expansion factor is set to 1.5; when the standard deviation is between 10 and 30, the expansion factor is set to 1.2; and when the standard deviation is greater than 30, the expansion factor is set to 0.8. Dynamic radius step calculation example: When the dynamic angle step is 10 degrees and the dynamic expansion factor is 1.2, the dynamic radius step is 12 pixels.

[0079] The initial scan coordinates are calculated based on the dynamic angle step and the dynamic radius step. Scanning is performed using polar coordinates, starting at 0 degrees and increasing by the dynamic angle step until it reaches 360 degrees. The radius also increases by the dynamic radius step, starting at the image center and continuing until it reaches the image boundary. For each angle and radius combination, the polar coordinates are converted to rectangular coordinates to obtain the initial scan coordinates. For example, if the center coordinates are (320, 240), the angle is 30 degrees, and the radius is 60 pixels, the corresponding initial scan coordinates are (320 + 60 × cos(30°), 240 + 60 × sin(30°)), or (372, 270).

[0080] The local response intensity value is calculated based on the initial scan coordinates. This value is obtained by calculating the grayscale gradient amplitude within a 5×5 pixel window surrounding the initial scan coordinates. The specific calculation method is to calculate the horizontal and vertical grayscale differences for each pixel in the window, then take the square root of the sum of the squares of these two differences as the gradient amplitude for that pixel. Finally, the average gradient amplitude of all pixels in the window is calculated as the local response intensity value. For example, the average gradient amplitude within a 5×5 window at a certain initial scan coordinate is 25.

[0081] The focus adjustment coefficient is calculated by dividing the local response strength by a preset baseline threshold. The preset baseline threshold can be set based on image characteristics; for example, it can be set to 20 for a typical scene image. For a local response strength of 25 and a preset baseline threshold of 20, the focus adjustment coefficient is 25 / 20 = 1.25. A focus adjustment coefficient greater than 1 indicates a region rich in detail and requiring more sampling; a focus adjustment coefficient less than 1 indicates a relatively flat region and suitable for sparse sampling.

[0082] The initial scan coordinates are sampled and adjusted based on the focus adjustment coefficient to generate the dynamic focus scan position. This adjustment is performed by offsetting the initial scan coordinates toward the image detail. The offset distance is equal to the product of the focus adjustment coefficient and the base offset distance. The base offset distance can be set to 5 pixels, and the detail direction is determined by calculating the local gradient direction. For example, when the focus adjustment coefficient is 1.25, the base offset distance is 5 pixels, and the local gradient direction is 45 degrees, the adjusted coordinate offset is (1.25 × 5 × cos(45°), 1.25 × 5 × sin(45°)), or (4.42, 4.42). If the initial scan coordinates are (372, 270), the adjusted dynamic focus scan position is (376, 274).

[0083] The first pixel window is selected based on the dynamic focus scan position. The size of the first pixel window can be set to 9×9 pixels, centered at the dynamic focus scan position. The local texture complexity is generated by calculating the cumulative difference between the pixels in the first pixel window and the weighted average of the window. The window weighted average is calculated using Gaussian weights, with the center pixel having the largest weight and gradually decreasing outward. The cumulative difference is the sum of the absolute values of the difference between the grayscale value of each pixel in the window and the weighted average. For example, if the weighted average of the 9×9 window at a dynamic focus scan position is 128, and the sum of the absolute values of the differences between all pixels in the window and the average is 560, then the local texture complexity is 560.

[0084] A second pixel window is selected based on the dynamic focus scan position. The second pixel window can be set to 15×15 pixels and is also centered at the dynamic focus scan position. The neighborhood grayscale uniformity is calculated by calculating the ratio of the pixel dispersion of the second pixel window to the overall image dispersion. The pixel dispersion is calculated using the standard deviation of the pixel grayscale values within the window. For example, if the standard deviation of the pixel grayscale values within a 15×15 window at a dynamic focus scan position is 15 and the standard deviation of the overall image pixel grayscale values is 30, then the neighborhood grayscale uniformity is 15 / 30 = 0.5.

[0085] Multiply the local texture complexity and the neighborhood grayscale uniformity by the corresponding preset weights to obtain the first weighted feature and the second weighted feature. The preset weight of the local texture complexity can be set to 0.6, and the preset weight of the neighborhood grayscale uniformity can be set to 0.4. When the local texture complexity is 560 and the neighborhood grayscale uniformity is 0.5, the first weighted feature is 560×0.6=336, and the second weighted feature is 0.5×0.4=0.2. The basic area feature value is determined to be the sum of the first weighted feature and the second weighted feature, that is, 336+0.2=336.2.

[0086] The boundary attenuation coefficient is calculated based on the distance from the dynamic focus scan position to the image boundary. As the scan position approaches the image boundary, the boundary attenuation coefficient gradually decreases, making the edge area less interesting. The boundary attenuation coefficient is calculated by taking the minimum distance from the dynamic focus scan position to the four boundaries and dividing it by one-quarter of the image diagonal length, with the result constrained to a range between 0.5 and 1. For example, for a 640×480 image, the diagonal length is approximately 800 pixels, and one-quarter is 200 pixels. If the distance from a dynamic focus scan position to the nearest boundary is 150 pixels, the boundary attenuation coefficient is min(max(150 / 200,0.5),1)=0.75.

[0087] The region's interest level is calculated by multiplying the base region's eigenvalue by the boundary attenuation coefficient. For the example above, the region's interest level is 336.2 × 0.75 = 252.15. The resulting region's interest level is used to assess the importance of an image region; a higher interest level indicates a more noteworthy region.

[0088] Traditional image region interest calculations mainly use fixed-step grid scanning and saliency-based region division methods. These methods use uniform parameters for images of different complexities, resulting in wasted computing resources or omission of important information. Existing adaptive sampling methods usually only consider local gradient features and ignore important information such as texture complexity and grayscale uniformity. The method of this embodiment realizes adaptive sampling in polar coordinates by introducing dynamic angle step and dynamic radius step, and combines the dynamic focus adjustment mechanism to make the sampling position more consistent with the local characteristics of the image; at the same time, it comprehensively considers local texture complexity and neighborhood grayscale uniformity, and introduces a boundary attenuation mechanism to make the regional interest calculation more comprehensive and accurate.

[0089] In an optional embodiment, adaptive region expansion is performed with the highest regional interest point as the center to obtain face positioning parameters, and extracting the target face image includes:

[0090] Determining a region growth starting point at a location with a maximum region interest, constructing a polar coordinate grid with the region growth starting point as the center, calculating an interest gradient value on each directional axis of the polar coordinate grid, multiplying the interest gradient value by the directional axis length to generate an axial expansion distance, determining boundary point coordinates based on the axial expansion distance, and performing a low-pass filter on the sequence of boundary point coordinates to obtain a smooth region outline;

[0091] Ellipse fitting is performed on the smooth area contour to calculate the area center coordinates, major axis length, minor axis length and direction angle, and combine them to generate a positioning parameter set. A coordinate transformation matrix is constructed according to the positioning parameter set. The coordinate transformation matrix is applied to the preprocessed image sequence for geometric mapping, and size normalization is performed to extract the target face image.

[0092] In one specific embodiment, preprocessing operations are performed on the original captured images, including image resizing, brightness normalization, and contrast enhancement, to generate a preprocessed image sequence. After preprocessing, a facial feature detection algorithm is used to calculate the regional interest value of each pixel in the image. A higher interest value indicates a greater probability that the pixel is located in the face region.

[0093] To locate the face region, the maximum position is searched in the regional interest distribution map. This position usually corresponds to the center of the face and is set as the starting point for region growth. The coordinates are marked as (x0, y0). A polar coordinate grid is constructed with this starting point as the center. In the embodiment, the polar coordinate grid consists of 16 directional axes, with directional angles uniformly distributed between 0 and 360 degrees. The angle of each directional axis can be expressed as θi = i × 22.5 degrees, where i ranges from 0 to 15.

[0094] On each directional axis, starting from the starting point, the interest gradient is calculated along the axial direction. The interest gradient is calculated as the difference between the interest values of two adjacent points. Assuming that the interest value at a point r away from the starting point in direction θi is I(r, θi), the interest gradient at that point can be expressed as the rate of change of the interest value. When the gradient value changes from positive to negative and the amplitude exceeds a preset threshold (set to 0.15 in this embodiment), it indicates that the face region boundary has been reached.

[0095] The gradient value on each directional axis is multiplied by the length of the corresponding directional axis to generate the axial expansion distance Di. In this embodiment, if a boundary feature is detected in direction θi at a distance from the starting point ri, the axial expansion distance Di in that direction = ri. For example, for an actual face image, if the distance to the boundary point detected in the 0-degree direction is 35 pixels, then D0 = 35; if the distance to the boundary point detected in the 22.5-degree direction is 38 pixels, then D1 = 38, and so on.

[0096] Based on the axial expansion distance, the coordinates of the boundary points in each direction are determined. The boundary point coordinates are calculated as: xi = x0 + Di × cos (θi), yi = y0 + Di × sin (θi). In this way, a series of coordinate points (xi, yi) are obtained to form the initial area contour. Due to the noise interference in the detection process, the initial contour is not smooth enough, and a low-pass filter is performed on the boundary point coordinate sequence. In the embodiment, a 5-point weighted average filter is used, and the filter coefficients are [0.1, 0.2, 0.4, 0.2, 0.1]. After filtering, a smooth area contour is obtained, which more accurately represents the face boundary.

[0097] Ellipse fitting is performed on the smooth region outline, and the least squares method is used to calculate the best-fit ellipse parameters. The ellipse fitting result includes the region center coordinates (xc, yc), the major axis length a, the minor axis length b, and the orientation angle α. In this example, for a typical face image, the fitting parameters are: xc=120, yc=150, a=45, b=35, and α=15 degrees, indicating that the face region center is located at (120, 150), the major axis length is 45 pixels, the minor axis length is 35 pixels, and the major axis makes an angle of 15 degrees with the horizontal direction.

[0098] These parameters are combined to generate the positioning parameter set P = {xc, yc, a, b, α}, which is used to construct the coordinate transformation matrix T. The coordinate transformation matrix includes translation, rotation, and scaling operations, which are used to transform the fitted ellipse area into a standard face image. The translation operation moves the center of the ellipse to the coordinate origin, the rotation operation aligns the ellipse's major axis with the coordinate axis, and the scaling operation adjusts the ellipse to a standard size.

[0099] Apply the coordinate transformation matrix T to the preprocessed image sequence for geometric mapping. For each pixel (x, y) in the preprocessed image, the inverse transformation T -1 Calculate the corresponding position in the original image and extract the pixel value at that position. This method can correct the tilt and size change of the face in the image.

[0100] Perform size normalization to resize the transformed image to a fixed size (64×64 pixels in this example) to obtain a standardized facial image. This normalization ensures that facial images collected under different conditions have the same size and pose, facilitating subsequent recognition algorithm processing.

[0101] In this embodiment, by determining the starting point of region growth at the maximum value of interest, the region to be segmented can be accurately located without human intervention, thereby improving processing efficiency and stability; the expansion distance is calculated by using a polar coordinate grid and the product of the interest gradient and the direction axis length, which can take into account the morphological characteristics of the region in different directions and achieve accurate detection of boundary points; low-pass filtering is performed on the sequence composed of boundary points to effectively suppress noise and small mutations, making the obtained regional contour smoother and more natural, and improving the subsequent fitting accuracy; ellipse fitting is performed based on the smooth contour to directly obtain positioning parameters such as the region center, major and minor axis lengths, and direction angles, thereby achieving an accurate geometric description of the target region; a coordinate transformation matrix is constructed using the positioning parameter set to perform geometric correction and size standardization on the preprocessed image sequence to ensure the consistency of the extracted facial images in position, scale, and direction, thereby facilitating subsequent recognition or analysis tasks.

[0102] In an optional embodiment, calculating phase spectrum change features for continuous target facial images, determining a state stability score in combination with stability prior information, and generating a shooting trigger signal based on an adaptive threshold comprises:

[0103] Performing a two-dimensional Fourier transform on continuous target facial images to obtain a frequency domain representation, extracting a phase component from the frequency domain representation to generate a phase spectrum sequence, and stacking the phase spectrum sequence in the time dimension to form a phase spectrum tensor;

[0104] Calculating the phase spectrum differences of adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence, constructing a frequency weight function based on a Gaussian kernel, and performing a weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain a phase spectrum change feature;

[0105] Acquire stability prior information of the face area, map the stability prior information to the frequency space to obtain a state evaluation parameter, and generate a state stability score according to the combined distribution law of the phase spectrum change characteristics and the state evaluation parameter;

[0106] Calculate the mean of the state stability scores within the historical time window, determine the dynamic adjustment coefficient by the ratio of the mean to the preset stability threshold, and adjust the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive judgment threshold; when the state stability scores are all greater than the adaptive judgment threshold within a continuous preset frame number threshold, a shooting trigger signal is generated.

[0107] In a specific embodiment, a two-dimensional Fourier transform operation is performed on the collected continuous facial image sequence to convert the image from the spatial domain to the frequency domain. Specifically, for each frame of the facial image in the time series, a two-dimensional fast Fourier transform (FFT) is performed to obtain a frequency domain representation containing an amplitude spectrum and a phase spectrum. The phase component is extracted from the frequency domain representation to form a phase spectrum sequence, and the phase spectrum contains image structure information and motion information. For example, for a facial image with a size of 256×256 pixels, a phase spectrum representation of the same size is obtained after the transformation. These phase spectra are stacked in chronological order to form a three-dimensional phase spectrum tensor. If 10 frames of images are collected continuously, the size of the phase spectrum tensor formed is 256×256×10.

[0108] For the resulting phase spectrum tensor, the phase spectrum differences between adjacent time frames are calculated. For example, for the phase spectra of frame t and frame t+1, the phase difference between the corresponding pixels is calculated to obtain the phase spectrum difference, which reflects the degree of change between the two frames. For n frames in a time series, n-1 phase spectrum difference results are obtained. A frequency weighting function based on a Gaussian kernel is then constructed. This function assigns different weights to different frequency components, with higher weights given to mid- and low-frequency regions (corresponding to the primary frequency characteristics of facial motion) and lower weights given to high-frequency regions. In practice, the standard deviation of the Gaussian kernel function can be set to 20% of the frequency domain size, that is, for a 256×256 frequency domain representation, the standard deviation is set to 51.2. The frequency weighting function is weighted and combined with the phase spectrum difference sequence to generate a phase spectrum change feature. This weighting operation enhances sensitivity to changes in facial motion while suppressing noise interference.

[0109] While acquiring phase spectrum variation features, prior stability information for the facial region is also obtained. This prior information includes facial posture angles, degree of expression change, blinking status, and more. For example, a head pitch angle and left-right rotation angle determined by feature point detection that do not exceed 15 degrees is considered posture stability; a displacement of key points of the mouth and eyebrows that is less than 5 pixels is considered expression stability; and a change in eye opening and closing that is less than 10% is considered blinking stability. This prior stability information is converted to frequency space using a mapping function to obtain state assessment parameters. The mapping process takes into account the frequency domain characteristics of different stability factors. For example, changes in head posture primarily affect low-frequency regions, while changes in expression correspond to mid-frequency regions. The state assessment parameters are represented as frequency domain distributions, sharing the same dimensional structure as the phase spectrum variation features.

[0110] A state stability score is calculated based on the combined distribution of phase spectrum characteristics and state assessment parameters. This combination is performed by taking the dot product of the two and then normalizing them. The score ranges from 0 to 1, with 1 indicating complete stability and 0 indicating extreme instability. In practical applications, a stable state score is typically above 0.85. This score is continuously calculated and updated to form a state stability score sequence.

[0111] To adapt to different scenarios and user states, an adaptive threshold judgment mechanism is implemented. The average state stability score is calculated within a historical time window, which can be set to 3 seconds (approximately 90 frames). The ratio of the calculated average to a preset stability threshold (e.g., 0.80) is used as the dynamic adjustment coefficient. When the user struggles to maintain high stability in certain environments (such as low light or motion), the average drops to 0.70, resulting in a dynamic adjustment coefficient of 0.875. Based on this, the preset threshold is adjusted downward, resulting in a more relaxed adaptive judgment threshold of 0.70. Conversely, under ideal conditions, when user stability improves, the adaptive threshold increases accordingly to ensure optimal shooting quality.

[0112] The relationship between the state stability score and the adaptive judgment threshold is continuously monitored. When the state stability score exceeds the adaptive judgment threshold for a preset frame number threshold (for example, 15 frames, corresponding to approximately 0.5 seconds), the system considers the target face to be stable and generates a capture trigger signal. This signal can trigger the camera to actually capture the subject or mark a keyframe in the video recording, thereby capturing the subject's portrait in the most stable state.

[0113] In this embodiment, phase spectrum difference is used to highlight subtle phase changes between consecutive frames, which can more sensitively capture small movements of facial expressions; frequency domain phase information is insensitive to amplitude noise and light fluctuations, making the state stability assessment more robust and less susceptible to environmental interference; the judgment threshold is dynamically adjusted based on the stability score of the historical window, which can adapt to different facial motion characteristics and scene changes, reducing false triggering or missed triggering; the stability prior information of the facial area is mapped to the frequency space, and combined with the phase spectrum change characteristics for joint evaluation, which helps to fuse the spatial and frequency domain priors and improve the reliability of trigger judgment; when multiple consecutive frames meet the adaptive threshold conditions, the shooting signal is immediately output to automatically capture the "optimal still state" moment, which is suitable for applications such as automatic photography and video still frame selection.

[0114] In an optional implementation, obtaining the stability prior information of the face region includes:

[0115] Obtaining coordinates of facial feature points in a target facial image, calculating displacement vectors of the facial feature point coordinates between adjacent image frames, and constructing a facial posture motion model based on the displacement vectors;

[0116] Performing an affine transformation on the facial feature point coordinates to obtain reference plane facial coordinates, calculating a feature distance matrix of the reference plane facial coordinates, and extracting facial region deformation parameters;

[0117] Calculating the translation component, rotation component and scaling component in the face posture motion model respectively, and combining them with the face region deformation parameter to obtain a face motion feature vector;

[0118] Performing principal component decomposition on the facial motion feature vector to obtain a feature projection matrix, projecting the facial motion feature vector onto the feature projection matrix to obtain a reduced-dimensional feature representation, and constructing a stability evaluation space based on the reduced-dimensional feature representation;

[0119] A local optimal stability interval is extracted from the stability evaluation space, and statistical features of the local optimal stability interval are output as stability priori information of the face region.

[0120] In a specific embodiment, the coordinates of facial feature points in the target face image are obtained. Specifically, a feature point detector based on deep learning can be used to extract key feature points in the target face image, including 68 feature points such as the corners of the eyes, the tip of the nose, and the corners of the mouth. For the input video sequence, face detection and feature point positioning are performed on each frame of the image to obtain a feature point coordinate set P = {p1, p2, ..., p 68}, where each feature point p i Contains two-dimensional coordinates (x, y).

[0121] Calculate the displacement vector of the facial feature point coordinates between adjacent image frames. For the feature point set P of the tth frame and the t+1th frame t and P t+1 , calculate the displacement vector D={d1, d2, ..., d 68}, where d i Represents the feature point p i The displacement from frame t to frame t+1. For example, if the coordinates of feature point p3 in two adjacent frames are (127, 185) and (130, 187), the displacement vector d3 = (3, 2).

[0122] The face posture motion model is constructed based on the displacement vector. The rigid transformation model is used to describe the changes in face posture. The displacement vectors of the feature points are fitted by the least squares method to obtain the rigid transformation matrix T containing translation, rotation and scaling parameters. The matrix can be decomposed into the translation vector (t x , t y ), rotation angle θ and scaling factor s. For example, the transformation parameters between two adjacent frames are: translation vector (2.5, 1.8) pixels, rotation angle 1.2 degrees, scaling factor 1.03.

[0123] Perform affine transformation on the facial feature point coordinates to obtain the reference plane facial coordinates. Select a standard frontal face pose as the reference plane, and map the facial feature points of each frame to the reference plane through affine transformation to eliminate the influence of pose changes. The transformed feature point set is recorded as P'={p'1, p'2, ..., p' 68}. Calculate the affine transformation matrix with the corners of the eyes and the tip of the nose as corresponding points, and map all feature points to the reference plane.

[0124] Calculate the feature distance matrix of the reference plane face coordinates. On the reference plane, calculate the Euclidean distance between each feature point to form a distance matrix M with a size of 68×68. Each element m in the matrix ij Represents the feature point p' i and p' j By comparing the changes in the corresponding distances between different frames, the degree of facial deformation can be quantified. For example, if the distance between the corners of the mouth changes from 52 pixels in frame t to 58 pixels in frame t+1, it indicates that the mouth has opened.

[0125] Extract deformation parameters of the facial region. Deformation parameters are extracted based on changes in the feature distance matrix, focusing on areas prone to deformation, such as the eyes, mouth, and eyebrows. The rate of change of distance between these regions between consecutive frames is calculated to construct a deformation parameter vector F. For example, for the mouth region, parameters such as the rate of change of distance between the upper and lower lips and the rate of change of distance between the corners of the mouth can be extracted to form a deformation parameter vector with a dimension of 15.

[0126] Calculate the translation component, rotation component and scaling component in the face posture motion model respectively. Decompose the transformation matrix T obtained above to obtain the translation vector (t x , t y ), rotation angle θ and scaling factor s. In order to better represent the motion trend, the average value and standard deviation of 5 consecutive frames are calculated to obtain the translation component vector (t x _mean,t x _std, t y _mean,t y _std), a rotation component vector (θ_mean, θ_std), and a scaling component vector (s_mean, s_std).

[0127] The components of the facial pose motion model are combined with the facial region deformation parameters to generate a facial motion feature vector. The translation component vector, rotation component vector, scaling component vector, and deformation parameter vector F are concatenated to form a complete facial motion feature vector V with a dimension of 23. This feature vector comprehensively expresses both rigid motion and non-rigid deformation information of the face.

[0128] Principal component decomposition (PCD) is performed on facial motion feature vectors to obtain a feature projection matrix. Facial motion feature vectors from a large number of video sequences are collected to construct a sample set. PCA is applied to this sample set to extract the main directions of change and obtain the feature projection matrix W. In practice, the first 10 principal components are selected to retain approximately 95% of the information.

[0129] The facial motion feature vectors are projected onto the feature projection matrix to obtain a reduced-dimensional feature representation. For each facial motion feature vector V, the projection matrix W is used to reduce the dimensionality, resulting in a 10-dimensional reduced-dimensional feature representation V'. This reduced-dimensional feature effectively captures the main patterns of facial motion and deformation.

[0130] A stability evaluation space is constructed based on the reduced-dimensionality feature representation. A stability metric function is defined within the reduced-dimensional feature space, and a stability score is calculated for each feature point in this space. The stability score takes into account factors such as the average amplitude of feature point displacement, directional consistency, and periodicity. For example, in a 15-frame video, the stability score of feature points in the corners of the eyes is 0.85 (out of a maximum score of 1), the tip of the nose is 0.92, and the corners of the mouth are 0.73.

[0131] The local optimal stability interval is extracted from the stability evaluation space. A sliding window method is used to scan the entire video sequence, with a window size of 20 frames and a step size of 5 frames. A comprehensive stability score is calculated for each window, and the window with the highest local score is selected as the optimal stability interval. For example, in a 100-frame video, two local optimal stability intervals are identified: frames 25-45 and frames 60-80, with stability scores of 0.88 and 0.91, respectively.

[0132] The statistical characteristics of the local optimal stability interval are output as prior stability information for the face region. Statistical features such as the mean, standard deviation, and curvature of the feature point positions within these intervals are extracted to form prior stability information. The output prior information includes the frame index of the stable interval, the average position of each feature point, a confidence score, and a recommendation for the optimal frame suitable for face recognition. This prior information can be used to improve algorithm performance in subsequent tasks such as face recognition, 3D reconstruction, and expression analysis.

[0133] Traditional facial stability assessment techniques primarily rely on image quality evaluation methods, such as static metrics like clarity, contrast, and brightness, or simple frame-to-frame difference calculations. These methods typically only consider image-level information and fail to fully utilize the motion information of facial landmarks for refined analysis. For example, typical quality-based methods use the Laplacian operator to calculate image clarity or histogram analysis to assess image brightness distribution, then select the frame with the highest quality as the optimal frame. Frame-to-frame difference-based methods typically calculate the sum of pixel differences between adjacent frames or the structural similarity (SSIM) metric, selecting the interval with the smallest change as the stable interval. These methods have limited effectiveness in assessing facial stability in complex scenarios, especially when the face undergoes changes in expression or slight posture, making it difficult to accurately assess the true stability of the face. The method of this embodiment elevates facial stability assessment from the image level to the feature point motion model level, and more accurately describes facial motion characteristics by constructing a facial posture motion model and analyzing non-rigid deformation parameters. A reference plane transformation mechanism is introduced to effectively separate rigid motion and non-rigid deformation, so that the stability of facial expressions can be accurately assessed even when the face posture changes. Principal component analysis is used for feature dimensionality reduction to construct a dedicated stability assessment space, reduce redundant information, and improve computational efficiency. A sliding window mechanism is designed to discover multiple local optimal stability intervals in video sequences, not just the global optimal solution, which is more in line with actual application needs.

[0134] Figure 3 The following is a performance comparison diagram of face stability assessment methods. The proposed method achieves significant advantages across all test scenarios. On the standard test set, the new method achieves 87.0% accuracy, outperforming traditional methods by 20.8 and frame difference methods by 15.0 percentage points. The advantages of the new method are particularly pronounced in complex scenarios, reaching 85.0% accuracy in pose variations, 83.0% in expression variations, and 81.0% in illumination variations.

[0135] Traditional quality assessment methods performed worst in scenes with changing facial expressions (only 56.0%), demonstrating their inability to cope with the challenges posed by non-rigid deformations. While the inter-frame difference method slightly outperformed traditional methods overall, it also performed poorly in scenes with changing facial expressions (60.0%). In contrast, the proposed method, based on a feature point motion model and dimensionality-reduced feature representation, is more effective in assessing facial stability and is particularly well-suited for identifying stable intervals in complex and changing scenes, providing reliable prior information for subsequent facial recognition and analysis tasks.

[0136] In an optional embodiment, multi-scale features of a facial image are extracted based on a pyramid decomposition structure, structural fingerprints and texture fingerprints are generated through frequency domain transform compression, and dual index features are formed by fusing them in a feature manifold space. The following steps are included:

[0137] A pyramid decomposition structure is constructed for standard face images. The edge direction histogram is extracted at each scale level of the pyramid decomposition structure. The edge direction histograms of adjacent scale levels are cross-compared to obtain a multi-scale structural feature matrix.

[0138] Calculating grayscale co-occurrence features at each scale level of the pyramid decomposition structure, accumulating the grayscale co-occurrence features in different directions to form a directional accumulation curve, and generating a multi-scale texture feature matrix based on the distribution law of the directional accumulation curve;

[0139] Performing discrete cosine transform on the multi-scale structural feature matrix to obtain a frequency domain coefficient matrix, selecting low-frequency components in the frequency domain coefficient matrix to construct a structural feature descriptor, and generating a structural fingerprint code based on the energy distribution of the structural feature descriptor;

[0140] Performing wavelet transform on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, extracting high-frequency components in the wavelet coefficient matrix to construct a texture response sequence, and generating a texture fingerprint code based on the energy distribution of the texture response sequence;

[0141] The structural fingerprint code and the texture fingerprint code are mapped to the local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and a dual index feature is obtained by adaptive fusion based on feature alignment loss optimization.

[0142] In one specific embodiment, a standard facial image is obtained and a pyramid decomposition structure is constructed for it. Specifically, the Gaussian pyramid decomposition method is used to sequentially downsample the original image to obtain multiple image levels of different resolutions. For a 1024×1024 pixel facial image, five scale levels, namely 512×512, 256×256, 128×128, and 64×64, can be obtained through continuous downsampling. At each scale level, the Sobel operator is applied to calculate the gradient magnitude and direction of the pixel points and extract the edge direction histogram. The edge direction interval is set to 0° to 180° and evenly divided into 12 bins, each with a width of 15°. The edge direction histogram of each scale level is normalized to ensure that the value range is between 0 and 1. The edge direction histograms of adjacent scale levels are cross-compared, and the difference values of each direction bin between each pair of adjacent levels are calculated to form a difference vector. These difference vectors are combined into a multi-scale structure feature matrix with a size of (number of levels - 1) × number of direction bins, that is, 4×12.

[0143] During the texture feature extraction stage, a gray-level co-occurrence matrix is calculated for each scale level of the pyramid decomposition structure, considering four directions (0°, 45°, 90°, and 135°), with a distance parameter set to 1 pixel. Statistical features such as contrast, energy, entropy, and correlation are extracted from the gray-level co-occurrence matrix. These statistical features are accumulated across different directions to generate a directional accumulation curve. For example, for the first scale level, the accumulated energy values for the 0° direction are 0.82, 0.75 for the 45° direction, 0.79 for the 90° direction, and 0.76 for the 135° direction. Based on the distribution pattern of the directional accumulation curve, the accumulated values of each statistical feature in different directions are sorted and combined to form a multi-scale texture feature matrix with dimensions of number of levels × number of features × number of directions, i.e., 5 × 4 × 4.

[0144] A discrete cosine transform is performed on the multi-scale structural feature matrix to obtain a frequency domain coefficient matrix. The low-frequency components in the upper left corner of the frequency domain coefficient matrix are selected to construct a structural feature descriptor, which reflects the primary energy distribution of facial structure. The energy values of these low-frequency components are calculated and binarized using a threshold of 0.5 to generate a structural fingerprint. Specifically, energy values greater than the threshold are assigned a value of 1, and otherwise 0. Taking the first scale level as an example, the extracted low-frequency component energy values are [0.87, 0.62, 0.43, 0.55, 0.32, 0.76, 0.41, 0.35, 0.29], corresponding to a structural fingerprint of [1, 1, 0, 1, 0, 1, 0, 0].

[0145] A wavelet transform is performed on the multi-scale texture feature matrix, and a three-level decomposition is performed using the Daubechies wavelet basis function to obtain a wavelet coefficient matrix. The coefficients of the high-frequency subbands (LH, HL, HH) are extracted from the wavelet coefficient matrix to construct a texture response sequence. For each scale level, three high-frequency subbands are obtained, and five statistics (mean, standard deviation, energy, entropy, and kurtosis) are extracted from each subband to form a 15-dimensional feature vector. Based on the energy distribution of the texture response sequence, a dynamic threshold is set for binarization processing to generate a texture fingerprint code. Taking the second scale level as an example, the energy distribution of the high-frequency subband LH is [0.23, 0.45, 0.67, 0.32, 0.51]. After applying a threshold of 0.4, the binary code [0, 1, 1, 0, 1] is obtained.

[0146] The structural fingerprint code and texture fingerprint code are mapped to the local feature manifold space, and the local linear embedding algorithm is used. The number of neighbor points is set to 10 and the embedding dimension is 20. A cross-modal correlation matrix is constructed in the local feature manifold space, and the matrix elements represent the similarity between the structural features and the texture features. The similarity is calculated using the cosine distance metric, with a value range of 0 to 1. For feature vectors of two different modalities, if the cosine similarity is greater than 0.75, they are considered to have a high correlation. Optimization is performed based on the feature alignment loss. The loss function combines the reconstruction error and the regularization term. The parameters are iteratively optimized using the gradient descent algorithm. The learning rate is set to 0.01 and the number of iterations is 500. By minimizing the alignment loss, the adaptive fusion of structural features and texture features is achieved, and a dual index feature with a dimension of 32 is obtained.

[0147] In this embodiment, the edge direction histogram and grayscale co-occurrence features are extracted at different scale levels through the pyramid decomposition structure, which can simultaneously capture details and global contour information, greatly improving the sensitivity and robustness to various scale features such as face deformation and expression changes; the structural fingerprint code reflects the overall shape and skeleton information of the face based on the low-frequency DCT component, and the texture fingerprint code characterizes the skin texture and details based on the high-frequency wavelet coefficient. The two complement each other, making the extracted features more discriminative and effectively improving the recognition accuracy; the grayscale co-occurrence matrix direction accumulation and edge direction histogram extraction both have statistical constraints on the local grayscale and edge layout, and can maintain stable feature responses under conditions of uneven lighting, weak texture or mild noise, thereby enhancing the system's anti-interference performance; the structural and texture fingerprint codes are mapped to the local feature manifold space, and by constructing a cross-modal correlation matrix and feature alignment loss optimization, the importance of the two types of features is adaptively balanced, and ultimately a more discriminative and generalizable dual index feature is obtained.

[0148] In an optional embodiment, the structural fingerprint code and the texture fingerprint code are mapped to a local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and a dual index feature is obtained by adaptive fusion based on feature alignment loss optimization, including:

[0149] Calculate the feature distance distribution of the structural fingerprint code and texture fingerprint code in the local neighborhood respectively, obtain the structural similarity matrix and texture similarity matrix, and construct the feature distribution representation;

[0150] Performing cross-correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross-modal correlation matrix, calculating feature weight coefficients, and performing weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix;

[0151] Constructing a nonlinear mapping model including an encoding layer and a decoding layer, inputting a structural fingerprint code and a texture fingerprint code, calculating the alignment loss between the mapping features based on the manifold consistency matrix, and constructing an alignment loss function in combination with the reconstruction loss;

[0152] Based on the alignment loss function, loss iterative optimization is performed to obtain an optimized nonlinear mapping function, and the optimized nonlinear mapping function is used to remap the structural fingerprint code and the texture fingerprint code to obtain the remapped feature;

[0153] The similarity between the remapped features is calculated to obtain a feature similarity matrix, feature fusion weights are calculated in combination with the feature distribution representation, and the remapped features are weightedly combined according to the feature fusion weights to obtain dual index features.

[0154] In a specific embodiment, the characteristic distance distribution of the structural fingerprint code and the texture fingerprint code in the local neighborhood is calculated respectively to obtain the structural similarity matrix and the texture similarity matrix, and construct the characteristic distribution representation. The structural fingerprint code is extracted from the key structural information of the face image, including the geometric relationship of organs such as the eyes, nose, and mouth, and the directional gradient histogram feature extraction method is used to obtain a 128-dimensional structural fingerprint code. The texture fingerprint code is extracted from the texture details of the face skin, and the local binary pattern algorithm is used to obtain a 256-dimensional texture feature vector. For each face image in the student photo, the local neighborhood range is defined as a 5×5 pixel area around the current feature point. In this neighborhood, the Euclidean distance between the structural fingerprint codes is calculated to construct a structural similarity matrix of size 128×128. Similarly, the Euclidean distance of the texture fingerprint code is calculated to obtain a 256×256 texture similarity matrix. For example, for the student photo with ID S20210523, the eigenvalues of the 15th and 16th dimensions of the extracted structural fingerprint are 0.82 and 0.76, respectively, with a Euclidean distance of 0.06, which are stored in the corresponding positions of the structural similarity matrix. These two matrices together constitute a feature distribution representation, describing the distribution relationship of facial features across different dimensions.

[0155] Cross-correlation analysis is performed on the local features of the structural and texture fingerprints to obtain a cross-modal correlation matrix. Feature weight coefficients are then calculated and weighted together with the feature distribution representation to obtain a manifold consistency matrix. Cross-correlation analysis is performed by calculating the Pearson correlation coefficient between the features of each dimension of the structural and texture fingerprints. For the student photo with ID S20210523, the correlation coefficient between the 20th dimension of the structural fingerprint and the 45th dimension of the texture fingerprint is 0.73, indicating a strong correlation between the features of these two dimensions. By calculating the correlation coefficients between all dimensions, a cross-modal correlation matrix of size 128×256 is constructed. Based on this matrix, the importance of each feature dimension is calculated to generate feature weight coefficients. The weight calculation uses the centrality algorithm, generating weight coefficients for each of the 128 dimensions of the structural fingerprint and the 256 dimensions of the texture fingerprint. For example, the weight of the 20th dimension of the structural fingerprint is 0.85, and the weight of the 45th dimension of the texture fingerprint is 0.78. These weight coefficients are weighted and combined with the previously obtained feature distribution representations (structural similarity matrix and texture similarity matrix) to obtain the manifold consistency matrix. This matrix describes the importance of different feature dimensions in preserving local geometric structure.

[0156] A nonlinear mapping model consisting of an encoding layer and a decoding layer was constructed. The structural and texture fingerprints were input. The alignment loss between the mapped features was calculated based on the manifold consistency matrix, and the alignment loss function was constructed by combining the reconstruction loss. The nonlinear mapping model uses a multilayer perceptron architecture and consists of two parts: encoding and decoding. The encoding part consists of three fully connected layers. For the structural fingerprint, the input layer has 128 neurons, and the hidden layers have 64 and 32 neurons, respectively. For the texture fingerprint, the input layer has 256 neurons, and the hidden layers have 128 and 32 neurons, respectively. The encodings of both fingerprints are ultimately mapped to a feature space of the same dimensionality (32). The decoding part, with a symmetrical structure, is used to reconstruct the mapped features back to the original feature space. For the photo of student S20210523, the structural fingerprint was encoded to produce a 32-dimensional mapped feature with values of 0.42 and 0.39 for the first and second dimensions, respectively. Similarly, the texture fingerprint had values of 0.38 and 0.41 for the first and second dimensions, respectively, after mapping. Based on the manifold consistency matrix, the alignment loss between the two mapped features is calculated, focusing on the degree of preservation of local geometric structure before and after mapping. A reconstruction loss is also calculated, representing the difference between the encoded and decoded features and the original features. The final alignment loss function is obtained by weighting the alignment loss with the reconstruction loss, with a weight of 0.7 for the alignment loss and a weight of 0.3 for the reconstruction loss.

[0157] Based on the alignment loss function, an iterative loss optimization is performed to obtain an optimized nonlinear mapping function. This optimized nonlinear mapping function is then used to remap the structural and texture fingerprints to obtain remapped features. The optimization process uses a stochastic gradient descent algorithm, with the learning rate initially set to 0.01 and halved every 5,000 iterations. To prevent overfitting, an L2 regularization term with a coefficient of 0.0001 is introduced. The model is trained on a student photo dataset with 1,000 batches, each containing 64 photos. After training, the optimized nonlinear mapping function is obtained. Taking the student photo with ID S20210523 as an example, the alignment loss before optimization was 0.385. After 50,000 iterations of optimization, the loss dropped to 0.092, demonstrating a significant improvement in the mapping effect. The optimized nonlinear mapping function is then used to remap the structural and texture fingerprints. After remapping, 32-dimensional remapped features are generated for the structural fingerprint and 32-dimensional remapped features for the texture fingerprint. For example, the first and second dimensions of the structural fingerprint code after remapping become 0.44 and 0.37 respectively, and the corresponding values of the texture fingerprint code after remapping are 0.43 and 0.36. It can be seen that the two features are closer in the mapping space.

[0158] The similarity between the remapped features is calculated to obtain a feature similarity matrix. Feature fusion weights are then calculated based on the feature distribution representation. The remapped features are then weighted according to the feature fusion weights to obtain dual-index features. The cosine similarity between the remapped features of the structural and texture fingerprints is calculated to generate a 32×32 feature similarity matrix. For student S20210523, the cosine similarity of the first dimension of the remapped structural and texture fingerprints is 0.95, indicating that these two dimensions are highly correlated in the new feature space. Based on the feature similarity matrix and the previously constructed feature distribution representation, an attention mechanism is used to calculate feature fusion weights. Specifically, a weight is assigned to each dimension based on the degree to which the local geometric structure is preserved in the feature distribution representation and its importance in the feature similarity matrix. The overall weight of the structural fingerprint remapped features is 0.52, and the overall weight of the texture fingerprint remapped features is 0.48. Each dimension also has a specific weight. For example, the weight of the first dimension of the structural fingerprint remapped features is 0.058, and the weight of the first dimension of the texture fingerprint remapped features is 0.053. According to these weights, the two remapped features are weighted and combined to produce a 32-dimensional dual-index feature. This feature integrates structural and texture information to more comprehensively characterize facial features and can be used for the retrieval and recognition of student photos. For student S20210523, the first dimension of the generated dual-index feature is 0.435. This is stored in the distributed database of the student management system to facilitate subsequent rapid retrieval and identity verification.

[0159] Table 1: Comparison of recognition accuracy of different feature representation methods:

[0160]

[0161] This solution demonstrated significant advantages under all test conditions. Under standard lighting conditions, it achieved a recognition accuracy of 96.8%, approximately 8 percentage points higher than traditional single-feature representation methods (87.4% for structural features and 89.2% for texture features), and 5.3 and 4.5 percentage points higher than simple feature concatenation and fixed-weight fusion methods, respectively.

[0162] This solution demonstrates exceptional robustness under complex conditions. Even under the challenging condition of a 30° posture deflection, it maintains an accuracy of 78.9%, compared to 62.4% and 59.3% for single-feature representation methods. In scenarios with changing facial expressions, this solution achieves an accuracy of 92.5%, significantly exceeding other methods.

[0163] Traditional facial feature extraction and student photo management technologies primarily employ single-feature representation methods, such as structural- or texture-based facial recognition, which fail to fully utilize multimodal facial information. Existing feature fusion methods typically employ simple feature concatenation or average fusion, failing to consider the correlation and complementarity between different features, resulting in poor fusion results. For example, some systems use principal component analysis to directly reduce the dimensionality before concatenating different features, or use fixed weights (e.g., 0.5:0.5) for feature weighting, ignoring the complex relationships between features. Furthermore, traditional methods exhibit poor robustness when dealing with nonlinear variations (such as changes in lighting, posture, and expression). The variable environment in which student photos are collected leads to reduced recognition accuracy.

[0164] The method of this embodiment proposes a dual feature representation of structural fingerprint code and texture fingerprint code to fully capture the structural and texture information of the face; introduces a manifold consistency matrix to describe the local geometric structure of different feature spaces, ensuring the consistency of local relations in the feature mapping process; designs a nonlinear mapping model based on alignment loss to align different modal features in a common feature space, enhancing the complementarity of features; adopts an adaptive feature fusion weight calculation method to dynamically adjust the fusion weight according to feature similarity and distribution characteristics, rather than simple averaging or splicing; implements a distributed storage and retrieval mechanism for student photos, improving the system's concurrent processing capability and retrieval efficiency; in order to solve the problems faced in student photo management, such as large differences in photo quality, variable collection environment, and insufficient recognition accuracy, especially in large-scale student data management scenarios, to improve recognition efficiency and accuracy. Through experimental verification, the method of this embodiment was tested on a student photo dataset containing 50,000 students. Compared with the traditional single feature method and the simple feature splicing method, the recognition accuracy was improved, while the feature dimension was reduced, significantly reducing storage and computing overhead. Under non-ideal conditions (such as lighting changes and posture deflection), the method of this embodiment is significantly more robust than the comparison method, and the decline in recognition accuracy is minimized. Furthermore, the use of distributed storage improves the system's concurrent processing capabilities and increases the speed of student photo retrieval, enabling it to support a student management system with tens of thousands of people online simultaneously. The method of this embodiment is not only applicable to student photo management but can also be extended to various facial recognition application scenarios, such as security monitoring and attendance management.

[0165] The embodiment of the present invention includes the following aspects:

[0166] The first unit is used to obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence;

[0167] The second unit is configured to perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood grayscale uniformity of the scanning position to obtain regional interest, perform adaptive region expansion with the highest point of regional interest as the center to obtain face positioning parameters, and extract the target face image;

[0168] The third unit is used to extract facial key points according to preset points based on the target face image, and calculate state evaluation parameters including facial posture angle value, eye opening value and expression parameter value based on the facial key points;

[0169] The fourth unit is used to calculate the phase spectrum change characteristics of continuous target face images, determine the state stability score based on the stability prior information, and generate a shooting trigger signal based on the adaptive threshold;

[0170] The fifth unit is used to respond to the shooting trigger signal, collect the face photo and enhance the processing to obtain the standard face image;

[0171] The sixth unit is used to extract multi-scale features of facial images based on the pyramid decomposition structure, generate structural fingerprints and texture fingerprints through frequency domain transformation compression, and fuse them in the feature manifold space to form dual index features;

[0172] The seventh unit is used to store the standard face image and the dual index features into the distributed storage system and generate an electronic certificate including the storage location.

[0173] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0174] processor;

[0175] a memory for storing processor-executable instructions;

[0176] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0177] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0178] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent collection and distributed storage management of student photos based on face recognition, characterized in that: include: Acquire a facial image video stream through a camera device, and preprocess it to obtain a preprocessed image sequence; Performing dynamic focus scanning on the preprocessed image sequence, calculating the local texture complexity and neighborhood grayscale uniformity of the scanning position to obtain the regional interest, including: A dynamic angle step is established based on the image center coordinates of the preprocessed image sequence, a dynamic radius step is generated by multiplying the dynamic angle step by the dynamic expansion factor, and the initial scanning coordinates are calculated based on the dynamic angle step and the dynamic radius step; Calculating a local response intensity value based on the initial scanning coordinates, generating a focus adjustment coefficient based on the ratio of the local response intensity value to a preset reference threshold, and adjusting the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position; A first pixel window is selected based on the dynamic focus scanning position, and the local texture complexity is generated by calculating the cumulative difference between the pixels in the first pixel window and the weighted average value of the window. A second pixel window is selected based on the dynamic focus scanning position, and the neighborhood grayscale uniformity is generated by calculating the ratio of the pixel discreteness of the second pixel window to the overall discreteness of the image. The local texture complexity and the neighborhood grayscale uniformity are multiplied by the corresponding preset weights to obtain the first weighted feature and the second weighted feature, and the basic area feature value is determined. The boundary attenuation coefficient is calculated according to the distance from the dynamic focus scanning position to the image boundary, and the product of the basic area feature value and the boundary attenuation coefficient is used to generate the regional interest degree; Adaptive region expansion is performed with the highest regional interest point as the center to obtain face positioning parameters and extract the target face image; Based on the target face image, facial key points are extracted according to preset points, and state evaluation parameters including facial posture angle value, eye opening value and expression parameter value are calculated based on the facial key points; The phase spectrum change characteristics of continuous target face images are calculated, the state stability score is determined based on the stability prior information, and the shooting trigger signal is generated based on the adaptive threshold. In response to the shooting trigger signal, the face photo is collected and enhanced to obtain a standard face image; The multi-scale features of facial images are extracted based on the pyramid decomposition structure, and the structural fingerprint and texture fingerprint are generated through frequency domain transformation compression, which are then fused in the feature manifold space to form a dual index feature. The standard face image and dual index features are stored in a distributed storage system to generate an electronic certificate containing the storage location.

2. The method according to claim 1, characterized in that Adaptive region expansion is performed with the highest regional interest point as the center to obtain face positioning parameters. Extracting the target face image includes: Determining a region growth starting point at a location with a maximum region interest, constructing a polar coordinate grid with the region growth starting point as the center, calculating an interest gradient value on each directional axis of the polar coordinate grid, multiplying the interest gradient value by the directional axis length to generate an axial expansion distance, determining boundary point coordinates based on the axial expansion distance, and performing a low-pass filter on the sequence of boundary point coordinates to obtain a smooth region outline; Ellipse fitting is performed on the smooth area contour to calculate the area center coordinates, major axis length, minor axis length and direction angle, and combine them to generate a positioning parameter set. A coordinate transformation matrix is constructed according to the positioning parameter set. The coordinate transformation matrix is applied to the preprocessed image sequence for geometric mapping, and size normalization is performed to extract the target face image.

3. The method according to claim 1, characterized in that Calculate the phase spectrum change characteristics of continuous target face images, combine the stability prior information to determine the state stability score, and generate the shooting trigger signal based on the adaptive threshold judgment, including: Performing a two-dimensional Fourier transform on continuous target facial images to obtain a frequency domain representation, extracting a phase component from the frequency domain representation to generate a phase spectrum sequence, and stacking the phase spectrum sequence in the time dimension to form a phase spectrum tensor; Calculating the phase spectrum differences of adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence, constructing a frequency weight function based on a Gaussian kernel, and performing a weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain a phase spectrum change feature; Acquire stability prior information of the face area, map the stability prior information to the frequency space to obtain a state evaluation parameter, and generate a state stability score according to the combined distribution law of the phase spectrum change characteristics and the state evaluation parameter; Calculate the mean of the state stability scores within the historical time window, determine the dynamic adjustment coefficient by the ratio of the mean to the preset stability threshold, and adjust the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive judgment threshold; when the state stability scores are all greater than the adaptive judgment threshold within a continuous preset frame number threshold, a shooting trigger signal is generated.

4. The method according to claim 3, characterized in that Obtaining stability prior information of the face area includes: Obtaining coordinates of facial feature points in a target facial image, calculating displacement vectors of the facial feature point coordinates between adjacent image frames, and constructing a facial posture motion model based on the displacement vectors; Performing an affine transformation on the facial feature point coordinates to obtain reference plane facial coordinates, calculating a feature distance matrix of the reference plane facial coordinates, and extracting facial region deformation parameters; Calculating the translation component, rotation component and scaling component in the face posture motion model respectively, and combining them with the face region deformation parameter to obtain a face motion feature vector; Performing principal component decomposition on the facial motion feature vector to obtain a feature projection matrix, projecting the facial motion feature vector onto the feature projection matrix to obtain a reduced-dimensional feature representation, and constructing a stability evaluation space based on the reduced-dimensional feature representation; A local optimal stability interval is extracted from the stability evaluation space, and statistical features of the local optimal stability interval are output as stability priori information of the face region.

5. The method according to claim 1, wherein The multi-scale features of facial images are extracted based on the pyramid decomposition structure. The structural fingerprint and texture fingerprint are generated through frequency domain transformation compression. The dual index features are fused in the feature manifold space, including: A pyramid decomposition structure is constructed for standard face images. The edge direction histogram is extracted at each scale level of the pyramid decomposition structure. The edge direction histograms of adjacent scale levels are cross-compared to obtain a multi-scale structural feature matrix. Calculating grayscale co-occurrence features at each scale level of the pyramid decomposition structure, accumulating the grayscale co-occurrence features in different directions to form a directional accumulation curve, and generating a multi-scale texture feature matrix based on the distribution law of the directional accumulation curve; Performing discrete cosine transform on the multi-scale structural feature matrix to obtain a frequency domain coefficient matrix, selecting low-frequency components in the frequency domain coefficient matrix to construct a structural feature descriptor, and generating a structural fingerprint code based on the energy distribution of the structural feature descriptor; Performing wavelet transform on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, extracting high-frequency components in the wavelet coefficient matrix to construct a texture response sequence, and generating a texture fingerprint code based on the energy distribution of the texture response sequence; The structural fingerprint code and the texture fingerprint code are mapped to the local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and a dual index feature is obtained by adaptive fusion based on feature alignment loss optimization.

6. The method according to claim 5, characterized in that The structural fingerprint code and the texture fingerprint code are mapped to the local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and the dual index features are obtained by adaptive fusion based on feature alignment loss optimization, including: Calculate the feature distance distribution of the structural fingerprint code and the texture fingerprint code in the local neighborhood respectively, obtain the structural similarity matrix and the texture similarity matrix, and construct the feature distribution representation; Performing cross-correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross-modal correlation matrix, calculating feature weight coefficients, and performing weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix; Constructing a nonlinear mapping model including an encoding layer and a decoding layer, inputting a structural fingerprint code and a texture fingerprint code, calculating the alignment loss between the mapping features based on the manifold consistency matrix, and constructing an alignment loss function in combination with the reconstruction loss; Based on the alignment loss function, loss iterative optimization is performed to obtain an optimized nonlinear mapping function, and the optimized nonlinear mapping function is used to remap the structural fingerprint code and the texture fingerprint code to obtain the remapped feature; The similarity between the remapped features is calculated to obtain a feature similarity matrix, feature fusion weights are calculated in combination with the feature distribution representation, and the remapped features are weightedly combined according to the feature fusion weights to obtain dual index features.

7. An intelligent collection and distributed storage management system for student photos based on face recognition, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence; The second unit is configured to perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood grayscale uniformity of the scanning position to obtain regional interest, perform adaptive region expansion with the highest point of regional interest as the center to obtain face positioning parameters, and extract the target face image; The third unit is used to extract facial key points according to preset points based on the target face image, and calculate state evaluation parameters including facial posture angle value, eye opening value and expression parameter value based on the facial key points; The fourth unit is used to calculate the phase spectrum change characteristics of continuous target face images, determine the state stability score based on the stability prior information, and generate a shooting trigger signal based on the adaptive threshold; The fifth unit is used to respond to the shooting trigger signal, collect the face photo and enhance the processing to obtain the standard face image; The sixth unit is used to extract multi-scale features of facial images based on the pyramid decomposition structure, generate structural fingerprints and texture fingerprints through frequency domain transformation compression, and fuse them in the feature manifold space to form dual index features; The seventh unit is used to store the standard face image and the dual index features into the distributed storage system and generate an electronic certificate including the storage location.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Face expression identification method based on video sequences

    CN105139004A

  • Human face image processing method and electronic device

    WO2021027585A1