Intelligent collection and distributed storage management method for school roll photos based on face recognition

Through dynamic focus scanning and adaptive region expansion technology, combined with facial key points and phase spectrum change characteristics, dual index features are generated for distributed storage, which solves the problems of inaccurate positioning and low efficiency in traditional student photo collection and storage, and achieves efficient and secure photo management.

CN120296185AActive Publication Date: 2025-07-11HANGZHOU RONGBO EDUCATION TECH CO LTD

Patent Information

Application Number
CN202510750475.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-11
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

传统学籍照片采集效率低下,照片质量不佳,定位不准确,存储管理效率低,难以支持大规模人脸图像的快速检索和比对。

Method used

Accurate face positioning is achieved through dynamic focus scanning and adaptive region expansion, real-time state evaluation is performed by combining facial key point extraction and phase spectrum change characteristics, standard face images are generated, and multi-scale features are extracted through pyramid decomposition to form dual index features for distributed storage.

Benefits of technology

It improves the accuracy and acquisition efficiency of face images, reduces the number of repeated shots, realizes efficient photo retrieval and security management, reduces storage costs and improves data access efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296185A_ABST
    Figure CN120296185A_ABST
Patent Text Reader

Abstract

The invention provides a school roll photo intelligent acquisition and distributed storage management method based on face recognition, which relates to the technical field of face recognition management, and comprises the following steps: acquiring a face image, executing dynamic focus scanning to position a face, extracting face key points and calculating state evaluation parameters; a shooting trigger signal is generated based on phase spectrum change features and stability prior information, a standard face image is acquired and enhanced, and a structure fingerprint code and a texture fingerprint code are extracted to form double index features. The intelligent degree of face photo collection is improved, the face recognition accuracy is enhanced, and efficient distributed storage management is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition management, and particularly to an intelligent acquisition and distributed storage management method for student status photos based on face recognition. Background Art

[0002] With the in-depth development of educational informatization construction, the photo acquisition and management in the student status management system have become an important part of school educational management work. Traditional student status photo acquisition usually adopts the methods of manual shooting, manual screening and centralized storage, which are inefficient in large-scale student status information management. In recent years, face recognition technology has been widely applied in various fields and has shown great value in identity verification, security monitoring and other aspects. Applying face recognition technology to the intelligent acquisition and management of student status photos can greatly improve the photo acquisition quality and management efficiency.

[0003] However, there are several problems in the acquisition, storage and management of student status photos. Traditional photo acquisition lacks an intelligent judgment mechanism and cannot evaluate the suitability of the face state in real time, resulting in problems such as closed eyes, tilted heads, improper expressions, etc. in the taken photos, and frequent retakes are required, wasting time and resources; the existing face localization algorithms have low accuracy under complex backgrounds or light conditions and are difficult to accurately extract a standardized face area, affecting subsequent processing and recognition effects; conventional photo storage methods often adopt a single centralized architecture, which not only has low retrieval efficiency, but also is prone to performance bottlenecks when the system load increases, and also lacks an effective feature indexing mechanism, making it difficult to support the fast retrieval and comparison requirements of large-scale face images.

[0004] With the increase in the number of students and the improvement of student status management requirements, there is an urgent need for a technical solution that can intelligently acquire high-quality photos and achieve efficient storage management to meet the needs of educational management. Summary of the Invention

[0005] An embodiment of the present invention provides an intelligent acquisition and distributed storage management method for student status photos based on face recognition, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiment of the present invention, an intelligent acquisition and distributed storage management method for student status photos based on face recognition is provided, including: Obtaining a face image video stream through a camera device and preprocessing it to obtain a preprocessed image sequence; Performing dynamic focus scanning on the preprocessed image sequence, calculating the local texture complexity and neighborhood gray level uniformity of the scanning position to obtain the region of interest, performing adaptive region expansion with the highest point of the region of interest as the center to obtain face localization parameters, and extracting the target face image; Based on the target face image, facial key points are extracted according to preset points, and state evaluation parameters including facial pose angle values, eye opening degree values, and expression parameter values are calculated based on the facial key points; Calculate the phase spectrum change features for consecutive target face images, determine the state stability score in combination with prior stability information, and judge and generate a shooting trigger signal based on an adaptive threshold; In response to the shooting trigger signal, collect a face photo and perform enhancement processing to obtain a standard face image; Extract multi-scale features of the face image based on the pyramid decomposition structure, generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation compression, and fuse them in the feature manifold space to form a dual-index feature; Store the standard face image and the dual-index feature in a distributed storage system, and generate an electronic voucher including the storage location.

[0007] In an alternative embodiment, perform dynamic focus scanning on the preprocessed image sequence, and calculate the local texture complexity and neighborhood gray uniformity of the scanning position to obtain the region of interest, including: Establish a dynamic angular step size based on the image center point coordinates of the preprocessed image sequence, multiply the dynamic angular step size by a dynamic expansion factor to generate a dynamic radius step size, and calculate the initial scanning coordinates based on the dynamic angular step size and the dynamic radius step size; Calculate the local response intensity value based on the initial scanning coordinates, generate a focus adjustment coefficient by taking the ratio of the local response intensity value to a preset reference threshold, and adjust the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position; Select a first pixel window based on the dynamic focus scanning position, generate the local texture complexity by calculating the cumulative difference between the pixel points in the first pixel window and the window weighted average value, select a second pixel window based on the dynamic focus scanning position, and generate the neighborhood gray uniformity by calculating the ratio of the pixel dispersion degree in the second pixel window to the overall image dispersion degree; Multiply the local texture complexity and the neighborhood gray uniformity by their corresponding preset weights respectively to obtain a first weighted feature and a second weighted feature, determine the basic region feature value, calculate the boundary attenuation coefficient according to the distance from the dynamic focus scanning position to the image boundary, and generate the region of interest by multiplying the basic region feature value by the boundary attenuation coefficient.

[0008] In an alternative embodiment, perform adaptive region expansion with the highest point of the region of interest as the center to obtain face positioning parameters, and extract the target face image, including: Determine the starting point of region growing at the position of the maximum regional interest degree. Construct a polar coordinate grid with the starting point of region growing as the center. Calculate the interest degree gradient value on each direction axis of the polar coordinate grid. Generate the axial expansion distance by multiplying the interest degree gradient value by the length of the direction axis. Determine the boundary point coordinates according to the axial expansion distance. Perform low-pass filtering on the sequence composed of the boundary point coordinates to obtain a smooth region contour; Perform elliptical fitting on the smooth region contour, calculate the region center coordinates, major axis length, minor axis length, and direction angle, and combine them to generate a set of positioning parameters. Construct a coordinate transformation matrix according to the set of positioning parameters. Apply the coordinate transformation matrix to the preprocessed image sequence for geometric mapping, and perform size normalization processing to extract the target face image.

[0009] In an alternative embodiment, calculating the phase spectrum change feature for consecutive target face images, combining the stability prior information to determine the state stability score, and judging and generating a shooting trigger signal based on an adaptive threshold includes: Perform two-dimensional Fourier transform on the consecutive target face images to obtain a frequency domain representation. Extract the phase components from the frequency domain representation to generate a phase spectrum sequence. Stack the phase spectrum sequence in the time dimension to form a phase spectrum tensor; Calculate the phase spectrum difference between adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence. Construct a frequency weight function based on a Gaussian kernel. Perform weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain the phase spectrum change feature; Obtain the stability prior information of the face region. Map the stability prior information to the frequency space to obtain a state evaluation parameter. Generate the state stability score according to the combined distribution law of the phase spectrum change feature and the state evaluation parameter; Calculate the mean value of the state stability scores within a historical time window. Determine the dynamic adjustment coefficient by the ratio of the mean value to a preset stability threshold. Adjust the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive decision threshold. Generate a shooting trigger signal when the state stability scores are all greater than the adaptive decision threshold within a consecutive preset frame number threshold.

[0010] In an alternative embodiment, obtaining the stability prior information of the face region includes: Obtain the coordinates of the face feature points in the target face image. Calculate the displacement vectors of the coordinates of the face feature points between adjacent image frames. Construct a face pose motion model based on the displacement vectors; Perform affine transformation on the coordinates of the face feature points to obtain the reference plane face coordinates. Calculate the feature distance matrix of the reference plane face coordinates. Extract the face region deformation parameters; Calculate the translation component, rotation component, and scaling component in the face pose motion model respectively, and combine them with the face region deformation parameters to obtain a face motion feature vector; Perform principal component analysis on the face motion feature vector to obtain a feature projection matrix, project the face motion feature vector onto the feature projection matrix to obtain a dimensionality-reduced feature representation, and construct a stability evaluation space based on the dimensionality-reduced feature representation; Extract the local optimal stable interval in the stability evaluation space, and output the statistical features of the local optimal stable interval as the stability prior information of the face region.

[0011] In an alternative embodiment, multi-scale features of a face image are extracted based on a pyramid decomposition structure, a structural fingerprint code and a texture fingerprint code are generated by frequency domain transformation compression, and a dual-index feature is formed by fusion in a feature manifold space, including: Construct a pyramid decomposition structure for a standard face image, extract the histogram of oriented gradients at each scale level of the pyramid decomposition structure, and perform cross-comparison on the histograms of oriented gradients at adjacent scale levels to obtain a multi-scale structural feature matrix; Calculate the gray-level co-occurrence features at each scale level of the pyramid decomposition structure, accumulate the gray-level co-occurrence features in different directions to form a direction accumulation curve, and generate a multi-scale texture feature matrix based on the distribution law of the direction accumulation curve; Perform discrete cosine transform on the multi-scale structural feature matrix to obtain a frequency domain coefficient matrix, select the low-frequency components in the frequency domain coefficient matrix to construct a structural feature descriptor, and generate a structural fingerprint code based on the energy distribution of the structural feature descriptor; Perform wavelet transform on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, extract the high-frequency components in the wavelet coefficient matrix to construct a texture response sequence, and generate a texture fingerprint code based on the energy distribution of the texture response sequence; Map the structural fingerprint code and the texture fingerprint code to a local feature manifold space, construct a cross-modal correlation matrix in the local feature manifold space, and perform adaptive fusion based on feature alignment loss optimization to obtain a dual-index feature.

[0012] In an alternative embodiment, mapping the structural fingerprint code and the texture fingerprint code to a local feature manifold space, constructing a cross-modal correlation matrix in the local feature manifold space, and performing adaptive fusion based on feature alignment loss optimization to obtain a dual-index feature includes: Calculate the feature distance distributions of the structural fingerprint code and the texture fingerprint code in the local neighborhood respectively, obtain a structural similarity matrix and a texture similarity matrix, and construct a feature distribution representation; Perform cross - correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross - modal correlation matrix, calculate the feature weight coefficients, and perform weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix; Construct a non - linear mapping model including an encoding layer and a decoding layer, input the structural fingerprint code and the texture fingerprint code, calculate the alignment loss between the mapping features based on the manifold consistency matrix, and construct an alignment loss function in combination with the reconstruction loss; Based on the alignment loss function, perform loss iterative optimization to obtain an optimized non - linear mapping function, and use the optimized non - linear mapping function to remap the structural fingerprint code and the texture fingerprint code to obtain remapped features; Calculate the similarity between the remapped features to obtain a feature similarity matrix, calculate the feature fusion weights in combination with the feature distribution representation, and perform weighted combination of the remapped features according to the feature fusion weights to obtain dual - index features.

[0013] In the second aspect of the embodiments of the present invention, there is provided an intelligent acquisition and distributed storage management system for student status photos based on face recognition, including: A first unit for obtaining a face image video stream through a camera device and pre - processing it to obtain a pre - processed image sequence; A second unit for performing dynamic focus scanning on the pre - processed image sequence, calculating the local texture complexity and neighborhood gray - scale uniformity of the scanning position to obtain the region of interest, performing adaptive region expansion with the highest point of the region of interest as the center to obtain face positioning parameters, and extracting the target face image; A third unit for extracting facial key points based on the target face image according to preset points, and calculating state evaluation parameters including facial pose angle values, eye opening degree values, and expression parameter values based on the facial key points; A fourth unit for calculating the phase spectrum change features of consecutive target face images, determining the state stability score in combination with the stability prior information, and judging and generating a shooting trigger signal based on an adaptive threshold; A fifth unit for responding to the shooting trigger signal, collecting a face photo and performing enhancement processing to obtain a standard face image; A sixth unit for extracting multi - scale features of the face image based on the pyramid decomposition structure, generating a structural fingerprint code and a texture fingerprint code through frequency - domain transformation compression, and fusing them in the feature manifold space to form dual - index features; A seventh unit for storing the standard face image and the dual - index features in a distributed storage system and generating an electronic voucher including the storage location.

[0014] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium storing computer program instructions thereon, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] In the embodiments of the present invention, the intelligent acquisition and distributed storage management method of student status photos based on face recognition realizes accurate face positioning through dynamic focus scanning and adaptive region expansion, effectively solves the problems of inaccurate positioning and susceptibility to environmental interference in traditional face acquisition, and improves the positioning accuracy and stability of face images; by extracting facial key points and calculating state evaluation parameters, combining the phase spectrum change characteristics and stability prior information, the real-time evaluation of the face state and the intelligent judgment of the best shooting time are realized, ensuring that the collected face photos meet the standard requirements, effectively reducing the number of manual interventions and repeated shootings, and improving the acquisition efficiency; by extracting structural features and texture features and compressing them to form dual-index features, combined with a distributed storage system, the efficient retrieval and secure management of student status photos are realized, not only reducing the storage cost, but also improving the data access efficiency and security, providing reliable technical support for the informatization management of student status. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of the intelligent acquisition and distributed storage management method of student status photos based on face recognition in the embodiments of the present invention; Figure 2 It is a schematic flowchart of dynamic focus scanning and region of interest calculation; Figure 3 It is a schematic diagram of the performance comparison of the face stability evaluation method. Detailed Embodiments

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0020] Figure 1 This is a schematic flowchart of the intelligent acquisition and distributed storage management method for student status photos based on face recognition according to an embodiment of the present invention. As Figure 1 shown, the method includes: Obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence; Perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood gray level uniformity of the scanning position to obtain the region of interest, perform adaptive region expansion with the highest point of the region of interest as the center to obtain face positioning parameters, and extract the target face image; Based on the target face image, extract facial key points according to preset points, and calculate state evaluation parameters including facial pose angle values, eye opening degrees, and expression parameter values based on the facial key points; Calculate the phase spectrum change characteristics of consecutive target face images, determine the state stability score in combination with prior stability information, and generate a shooting trigger signal based on an adaptive threshold; Respond to the shooting trigger signal, collect a face photo and perform enhancement processing to obtain a standard face image; Extract multi-scale features of the face image based on the pyramid decomposition structure, generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation compression, and fuse them in the feature manifold space to form a dual-index feature; Store the standard face image and the dual-index feature in a distributed storage system, and generate an electronic voucher including the storage location.

[0021] In an optional implementation manner, performing dynamic focus scanning on the preprocessed image sequence and calculating the local texture complexity and neighborhood gray level uniformity of the scanning position to obtain the region of interest includes: Establish a dynamic angle step based on the image center point coordinates of the preprocessed image sequence, multiply the dynamic angle step by a dynamic expansion factor to generate a dynamic radius step, and calculate the initial scanning coordinates according to the dynamic angle step and the dynamic radius step; Calculate the local response intensity value based on the initial scanning coordinates, generate a focus adjustment coefficient by taking the ratio of the local response intensity value to a preset reference threshold, and adjust the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position; Select a first pixel window based on the dynamic focus scanning position, generate the local texture complexity by calculating the cumulative difference between the pixel points in the first pixel window and the weighted average value of the window, select a second pixel window based on the dynamic focus scanning position, and generate the neighborhood gray level uniformity by calculating the ratio of the pixel dispersion degree in the second pixel window to the overall dispersion degree of the image; Multiply the local texture complexity and the neighborhood gray - level uniformity by their corresponding preset weights respectively to obtain the first weighted feature and the second weighted feature, determine the basic region eigenvalue, calculate the boundary attenuation coefficient according to the distance from the dynamic focus scanning position to the image boundary, and generate the region interest degree by multiplying the basic region eigenvalue by the boundary attenuation coefficient.

[0022] Figure 2 It is a schematic diagram of the dynamic focus scanning and region interest degree calculation process. As Figure 2 shown, in a specific embodiment, after pre - processing the image sequence, a dynamic angular step is established based on the image center - point coordinates. Specifically, the pre - processing of the image sequence includes image noise reduction, contrast adjustment, and color - space conversion, converting the RGB image into a grayscale image and removing noise through Gaussian filtering. The image center - point coordinates can be set at half of the image width and height. For example, for an image with a resolution of 640×480, the center - point coordinates are (320, 240). The dynamic angular step is adaptively adjusted according to the image complexity. When the overall image complexity is high, the angular step is set smaller, such as 5 degrees; when the overall image complexity is low, the angular step is set larger, such as 15 degrees. The image complexity is quantified by calculating the gray - level co - occurrence matrix features of the image, and the result value ranges from 0 to 1, where the larger the value, the higher the image complexity.

[0023] Multiply the dynamic angular step by the dynamic expansion factor to generate the dynamic radius step. The dynamic expansion factor is determined based on the image - region brightness distribution. The expansion factor is larger in regions with more uniform brightness, such as 1.5; the expansion factor is smaller in regions with drastic brightness changes, such as 0.8. In practical applications, the uniformity of the brightness distribution can be determined by calculating the brightness standard deviation of the local region of the image. When the standard deviation is less than 10, the expansion factor is set to 1.5; when the standard deviation is between 10 and 30, the expansion factor is set to 1.2; when the standard deviation is greater than 30, the expansion factor is set to 0.8. Example of calculating the dynamic radius step: When the dynamic angular step is 10 degrees and the dynamic expansion factor is 1.2, the dynamic radius step is 12 pixels.

[0024] Calculate the initial scanning coordinates according to the dynamic angular step and the dynamic radius step. Taking the image center - point as the origin, scanning is carried out in polar - coordinate mode. The angle starts from 0 degrees and increases in accordance with the dynamic angular step until 360 degrees; the radius starts from the image center - point and increases in accordance with the dynamic radius step until it reaches the image boundary. For each combination of angle and radius, convert the polar coordinates to rectangular coordinates to obtain the initial scanning coordinate points. For example, if the center - point coordinates are (320, 240), the angle is 30 degrees, and the radius is 60 pixels, then the corresponding initial scanning coordinate is (320 + 60×cos(30°), 240 + 60×sin(30°)), that is, (372, 270).

[0025] Calculate the local response intensity value based on the initial scan coordinates. The local response intensity value is obtained by calculating the magnitude of the gray-scale gradient within a 5×5 pixel window around the initial scan coordinates. The specific calculation method is to calculate the gray-scale differences in the horizontal and vertical directions for each pixel point within the window, then take the square root of the sum of the squares of the differences in the two directions as the gradient magnitude of the pixel point, and finally calculate the average value of the gradient magnitudes of all pixel points within the window as the local response intensity value. For example, the calculated average value of the gradient magnitudes within a 5×5 window at a certain initial scan coordinate is 25.

[0026] Generate a focus adjustment coefficient based on the ratio of the local response intensity value to a preset reference threshold. The preset reference threshold can be set according to the image characteristics. For example, for ordinary scene images, it can be set to 20. When the local response intensity value is 25 and the preset reference threshold is 20, the focus adjustment coefficient is 25 / 20 = 1.25. A focus adjustment coefficient greater than 1 indicates that the details in this area are rich and denser sampling is required; less than 1 indicates that this area is relatively flat and sparse sampling can be used.

[0027] Adjust the sampling position of the initial scan coordinates according to the focus adjustment coefficient to generate a dynamic focus scan position. The adjustment method is to offset the initial scan coordinates towards the image detail direction by a distance equal to the product of the focus adjustment coefficient and the base offset distance. The base offset distance can be set to 5 pixels, and the detail direction is determined by calculating the local gradient direction. For example, when the focus adjustment coefficient is 1.25, the base offset distance is 5 pixels, and the local gradient direction is 45 degrees, the offset amount of the adjusted coordinates is (1.25×5×cos(45°), 1.25×5×sin(45°)), that is, (4.42, 4.42). If the initial scan coordinates are (372, 270), then the adjusted dynamic focus scan position is (376, 274).

[0028] Select a first pixel window based on the dynamic focus scan position. The size of the first pixel window can be set to 9×9 pixels, centered on the dynamic focus scan position. Generate the local texture complexity by calculating the cumulative difference between the pixel points within the first pixel window and the weighted average of the window. The window weighted average is calculated using Gaussian weights, with the maximum weight at the center pixel and gradually decreasing towards the outside. The cumulative difference is the sum of the absolute values of the differences between the gray-scale values of each pixel point within the window and the weighted average. For example, the weighted average of a 9×9 window at a certain dynamic focus scan position is 128, and the sum of the absolute values of the differences between all pixel points within the window and this average is 560, then the local texture complexity is 560.

[0029] Select the second pixel window based on the dynamic focus scanning position. The second pixel window can be set to 15×15 pixels and is also centered on the dynamic focus scanning position. Calculate the neighborhood gray uniformity by computing the ratio of the pixel dispersion degree of the second pixel window to the overall image dispersion degree. The pixel dispersion degree is calculated using the standard deviation of the pixel gray values within the window. For example, if the standard deviation of the pixel gray values within a 15×15 window at a certain dynamic focus scanning position is 15 and the standard deviation of the overall image pixel gray values is 30, then the neighborhood gray uniformity is 15 / 30 = 0.5.

[0030] Multiply the local texture complexity and the neighborhood gray uniformity by their respective preset weights to obtain the first weighted feature and the second weighted feature. The preset weight for the local texture complexity can be set to 0.6, and the preset weight for the neighborhood gray uniformity can be set to 0.4. When the local texture complexity is 560 and the neighborhood gray uniformity is 0.5, the first weighted feature is 560×0.6 = 336, and the second weighted feature is 0.5×0.4 = 0.2. Determine the basic region feature value as the sum of the first weighted feature and the second weighted feature, that is, 336 + 0.2 = 336.2.

[0031] Calculate the boundary attenuation coefficient based on the distance from the dynamic focus scanning position to the image boundary. When the scanning position approaches the image boundary, the boundary attenuation coefficient gradually decreases, reducing the interest degree of the edge region. The boundary attenuation coefficient calculation method is: take the minimum distance from the dynamic focus scanning position to the four boundaries, divide it by one-fourth of the image diagonal length, and limit the result between 0.5 and 1. For example, for an image of 640×480, the diagonal length is approximately 800 pixels, and one-fourth is 200 pixels. If the distance from a certain dynamic focus scanning position to the nearest boundary is 150 pixels, then the boundary attenuation coefficient is min(max(150 / 200, 0.5), 1) = 0.75.

[0032] Generate the region interest degree by multiplying the basic region feature value by the boundary attenuation coefficient. For the above example, the region interest degree is 336.2×0.75 = 252.15. The finally obtained region interest degree is used to evaluate the importance of the image region, and the higher the interest degree, the more worthy of attention the region is.

[0033] Traditional methods for calculating the interest degree of image regions mainly use fixed-step grid scanning and saliency-based region division methods. These methods use unified parameters for images of different complexities, resulting in waste of computing resources or omission of important information. Existing adaptive sampling methods usually only consider local gradient features and ignore important information such as texture complexity and gray-level uniformity. The method of this embodiment realizes adaptive sampling in polar coordinates by introducing a dynamic angular step and a dynamic radius step. Combined with a dynamic focus adjustment mechanism, the sampling positions are more in line with the local characteristics of the image. At the same time, local texture complexity and neighborhood gray-level uniformity are comprehensively considered, and a boundary attenuation mechanism is introduced to make the calculation of region interest degree more comprehensive and accurate.

[0034] In an optional implementation manner, adaptive region expansion is performed with the highest point of region interest degree as the center to obtain face localization parameters. Extracting the target face image includes: Determine the starting point of region growth at the position of the maximum value of region interest degree. Construct a polar coordinate grid with the starting point of region growth as the center. Calculate the interest degree gradient value on each direction axis of the polar coordinate grid. Generate the axial expansion distance by multiplying the interest degree gradient value by the length of the direction axis. Determine the boundary point coordinates according to the axial expansion distance. Perform low-pass filtering on the sequence composed of the boundary point coordinates to obtain a smooth region contour; Perform ellipse fitting on the smooth region contour, calculate the region center coordinates, major axis length, minor axis length, and direction angle, and combine them to generate a set of localization parameters. Construct a coordinate transformation matrix according to the set of localization parameters. Apply the coordinate transformation matrix to the preprocessed image sequence for geometric mapping, and perform size normalization processing to extract the target face image.

[0035] In a specific implementation manner, preprocessing operations are performed on the originally acquired images, including image size adjustment, brightness normalization, and contrast enhancement, to generate a preprocessed image sequence. After preprocessing, use a face feature detection algorithm to calculate the region interest degree value of each pixel point in the image. The higher the interest degree value, the greater the probability that the point is located in the face region.

[0036] To locate the face region, search for the maximum value position in the region interest degree distribution map. This position usually corresponds to the center region of the face, and set it as the starting point of region growth, with the coordinates recorded as (x0, y0). Construct a polar coordinate grid with this starting point as the center. In the embodiment, the polar coordinate grid consists of 16 direction axes, and the direction angles are evenly distributed between 0 and 360 degrees. The angle of each direction axis can be expressed as θi = i × 22.5 degrees, where i ranges from 0 to 15.

[0037] On each direction axis, starting from the starting point, calculate the interest gradient value along the axis. The interest gradient value is calculated by the difference in interest between adjacent points. Assuming that the interest value at a distance r from the starting point in the direction θi is I(r, θi), the interest gradient value at this point can be expressed as the rate of change of the interest value. When the gradient value changes from positive to negative and its magnitude exceeds a preset threshold (set to 0.15 in the embodiment), it indicates that the boundary of the face region has been reached.

[0038] Multiply the gradient value on each direction axis by the corresponding axis length to generate the axial expansion distance Di. In the embodiment, if a boundary feature is detected at a distance ri from the starting point in the direction θi, the axial expansion distance Di in this direction is Di = ri. For example, for an actual face image, the distance of the boundary point detected in the 0-degree direction is 35 pixels, so D0 = 35; the distance of the boundary point detected in the 22.5-degree direction is 38 pixels, so D1 = 38, and so on.

[0039] Based on the axial expansion distance, determine the boundary point coordinates in each direction. The boundary point coordinates are calculated as: xi = x0 + Di × cos(θi), yi = y0 + Di × sin(θi). In this way, a series of coordinate points (xi, yi) are obtained, which constitute the initial region contour. Due to noise interference in the detection process, the initial contour is not smooth enough, so a low-pass filter is performed on the boundary point coordinate sequence. In the embodiment, 5-point weighted average filtering is used, and the filter coefficients are [0.1, 0.2, 0.4, 0.2, 0.1]. After filtering, a smooth region contour is obtained, which more accurately represents the face boundary.

[0040] Perform ellipse fitting on the smooth region contour, and use the least squares method to calculate the best-fitting ellipse parameters. The ellipse fitting results include the region center coordinates (xc, yc), the major axis length a, the minor axis length b, and the direction angle α. In the example, for a typical face image, the fitted parameters are: xc = 120, yc = 150, a = 45, b = 35, α = 15 degrees, indicating that the center of the face region is located at (120, 150), the major axis is 45 pixels long, the minor axis is 35 pixels long, and the major axis makes an angle of 15 degrees with the horizontal direction.

[0041] These parameters are combined to generate a set of positioning parameters P = {xc, yc, a, b, α}, which is used to construct a coordinate transformation matrix T. The coordinate transformation matrix includes translation, rotation, and scaling operations, and is used to transform the fitted ellipse region into a standard face image. The translation operation moves the ellipse center to the coordinate origin, the rotation operation aligns the major axis of the ellipse with the coordinate axes, and the scaling operation adjusts the ellipse to a standard size.

[0042] Apply the coordinate transformation matrix T to the preprocessed image sequence for geometric mapping. For each pixel point (x, y) in the preprocessed image, through the inverse transformation T -1Calculate its corresponding position in the original image and extract the pixel value at that position. This method can correct the tilt and size changes of the face in the image.

[0043] Perform size normalization processing to adjust the transformed image to a fixed size (64×64 pixels in the embodiment), obtaining a normalized face image. The normalization processing ensures that face images collected under different conditions have the same size and pose, facilitating subsequent processing by recognition algorithms.

[0044] In this embodiment, by determining the starting point of region growing at the maximum of the interest degree, the region to be segmented can be accurately located without manual intervention, improving the processing efficiency and stability; using the product of the polar coordinate grid and the interest degree gradient and the length of the direction axis to calculate the expansion distance can take into account the morphological characteristics of different directions of the region and achieve accurate detection of boundary points; performing low-pass filtering on the sequence composed of boundary points can effectively suppress noise and small mutations, making the obtained region contour smoother and more natural and improving the subsequent fitting accuracy; performing ellipse fitting based on the smoothed contour to directly obtain positioning parameters such as the region center, the lengths of the major and minor axes, and the direction angle, realizing accurate geometric description of the target region; using the set of positioning parameters to construct a coordinate transformation matrix to perform geometric correction and size normalization on the preprocessed image sequence, ensuring the consistency of the extracted face images in terms of position, scale, and direction, facilitating subsequent recognition or analysis tasks.

[0045] In an alternative embodiment, calculating the phase spectrum change feature for consecutive target face images, combining the prior information of stability to determine the state stability score, and judging and generating a shooting trigger signal based on an adaptive threshold includes: Performing two-dimensional Fourier transform on consecutive target face images to obtain a frequency domain representation, extracting the phase component from the frequency domain representation to generate a phase spectrum sequence, and stacking the phase spectrum sequence in the time dimension to form a phase spectrum tensor; Calculating the phase spectrum difference between adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence, constructing a frequency weight function based on a Gaussian kernel, and performing weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain the phase spectrum change feature; Obtaining the prior information of stability of the face region, mapping the prior information of stability to the frequency space to obtain a state evaluation parameter, and generating a state stability score according to the combined distribution law of the phase spectrum change feature and the state evaluation parameter; Calculating the mean value of the state stability scores within a historical time window, determining a dynamic adjustment coefficient as the ratio of the mean value to a preset stability threshold, adjusting the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive determination threshold; generating a shooting trigger signal when the state stability scores are all greater than the adaptive determination threshold within a consecutive preset number of frames threshold.

[0046] In a specific embodiment, a two-dimensional Fourier transform operation is performed on the collected continuous sequence of face images to transform the images from the spatial domain to the frequency domain. Specifically, for each frame of face image in the time series, a two-dimensional fast Fourier transform (FFT) is performed to obtain a frequency domain representation containing the amplitude spectrum and the phase spectrum. The phase component is extracted from the frequency domain representation to form a sequence of phase spectra, which contains the image structure information and the motion information. For example, for a face image with a size of 256×256 pixels, a phase spectrum representation of the same size is obtained after transformation. These phase spectra are stacked in chronological order to form a three-dimensional phase spectrum tensor. If 10 frames of images are continuously collected, the size of the formed phase spectrum tensor is 256×256×10.

[0047] For the formed phase spectrum tensor, calculate the phase spectrum difference between adjacent time frames. For example, for the phase spectra of the t-th frame and the (t + 1)-th frame, calculate the phase difference of the corresponding position pixels to obtain the phase spectrum difference reflecting the degree of change between the two frames. For n frames of images in the time series, n - 1 phase spectrum difference results can be obtained. At this time, construct a frequency weight function based on a Gaussian kernel, which assigns different weights to different frequency components. A higher weight is assigned to the middle and low frequency regions (corresponding to the main frequency characteristics of face motion), and a lower weight is assigned to the high frequency region. In practice, the standard deviation of the Gaussian kernel function can be set to 20% of the frequency domain size, that is, for a frequency domain representation of 256×256, the standard deviation is set to 51.2. The frequency weight function is weighted and combined with the phase spectrum difference sequence to generate the phase spectrum change feature. This weighting operation enhances the sensitivity to changes in the face motion state while suppressing noise interference.

[0048] While obtaining the phase spectrum change feature, obtain the prior information on the stability of the face region. These prior information include the pose angle of the face, the degree of expression change, the blinking state, etc. For example, the pitch angle and yaw angle of the head determined by feature point detection not exceeding 15 degrees can be regarded as pose stable; the displacement of the key points of the mouth and eyebrows less than 5 pixels can be regarded as expression stable; the change in the eye opening degree less than 10% can be regarded as blinking state stable. These stability prior information are transformed into the frequency space through a mapping function to obtain the state evaluation parameters. The mapping process takes into account the performance characteristics of different stability factors in the frequency domain. For example, the change in head pose mainly affects the low frequency region, while the expression change corresponds to the middle frequency region. The state evaluation parameters are represented in the form of a frequency domain distribution and have the same dimensional structure as the phase spectrum change feature.

[0049] Calculate the state stability score based on the combined distribution law of the phase spectrum change characteristics and the state evaluation parameters. The combination method is to perform dot multiplication of the two and then normalize the result. The score range is between 0 and 1, where 1 represents completely stable and 0 represents extremely unstable. In practical applications, the score in the stable state is usually above 0.85. Continuously calculate and update this score to form a state stability score sequence.

[0050] To adapt to the changes in different scenarios and user states, an adaptive threshold judgment mechanism is adopted. Calculate the mean value of the state stability scores within the historical time window, and the length of the time window can be set to 3 seconds (about 90 frames). Determine the dynamic adjustment coefficient as the ratio of the calculated mean value to the preset stability threshold (such as 0.80). When it is difficult for the user to maintain a high degree of stability in a specific environment (such as a low-light environment or a motion scenario), and the mean value drops to 0.70, the dynamic adjustment coefficient is 0.875 at this time. Accordingly, the preset threshold is lowered to obtain a more relaxed adaptive judgment threshold of 0.70. On the contrary, when the user's stability improves under ideal conditions, the adaptive threshold will be increased accordingly to ensure the best shooting quality.

[0051] Continuously monitor the relationship between the state stability score and the adaptive judgment threshold. When the state stability score is greater than the adaptive judgment threshold within the continuous preset number of frames threshold (for example, 15 frames, corresponding to about 0.5 seconds), the system considers that the target face is in a stable state, and then generates a shooting trigger signal. This signal can trigger the camera to perform an actual shooting operation, or mark key frames during video recording, so as to capture portraits in the best stable state.

[0052] In this embodiment, using the phase spectrum difference to highlight the subtle phase changes between consecutive frames can capture the tiny movements of facial expressions more sensitively; the frequency-domain phase information is insensitive to amplitude noise and light fluctuations, making the state stability evaluation more robust and less susceptible to environmental interference; dynamically adjusting the judgment threshold based on the stability score of the historical window can adapt to different facial motion characteristics and scene changes, reducing false triggers or missed triggers; mapping the prior information of the stability of the face region to the frequency space and jointly evaluating it in combination with the phase spectrum change characteristics helps to fuse the spatial and frequency-domain priors and improve the reliability of the trigger judgment; immediately outputting a shooting signal when multiple consecutive frames meet the adaptive threshold condition realizes the automatic capture of the "best static state" moment, which is applicable to applications such as automatic photography and video still frame selection.

[0053] In an alternative embodiment, obtaining the prior information of the stability of the face region includes: Obtain the coordinates of the facial feature points in the target face image, calculate the displacement vector of the facial feature point coordinates between adjacent image frames, and construct a facial pose motion model based on the displacement vector; Perform an affine transformation on the coordinates of the facial feature points to obtain the facial coordinates on the reference plane, calculate the feature distance matrix of the facial coordinates on the reference plane, and extract the deformation parameters of the facial region; Calculate the translation component, rotation component, and scaling component in the facial pose motion model respectively, and combine them with the deformation parameters of the facial region to obtain the facial motion feature vector; Perform principal component analysis on the facial motion feature vector to obtain the feature projection matrix, project the facial motion feature vector onto the feature projection matrix to obtain the dimensionality-reduced feature representation, and construct a stability evaluation space based on the dimensionality-reduced feature representation; Extract the locally optimal stable interval in the stability evaluation space, and output the statistical features of the locally optimal stable interval as the stability prior information of the facial region.

[0054] In a specific embodiment, obtain the coordinates of the facial feature points in the target facial image. Specifically, a feature point detector based on deep learning can be used to extract the key feature points in the target facial image, including 68 feature points such as the corners of the eyes, the tip of the nose, and the corners of the mouth. For the input video sequence, perform face detection and feature point localization on each frame of the image to obtain the feature point coordinate set P = {p1, p2,..., p 68}, where each feature point p i contains two-dimensional coordinates (x, y).

[0055] Calculate the displacement vector of the facial feature point coordinates between adjacent image frames. For the feature point sets P t and P t+1 of the t-th frame and the (t + 1)-th frame, calculate the displacement vector D = {d1, d2,..., d 68} of each corresponding feature point, where d i represents the displacement of the feature point p i from the t-th frame to the (t + 1)-th frame. For example, for the coordinates of the feature point p3 in two adjacent frames being (127, 185) and (130, 187) respectively, the displacement vector d3 = (3, 2).

[0056] Construct a facial pose motion model based on the displacement vector. Use a rigid transformation model to describe the change of the facial pose, and fit the displacement vector of the feature points by the least squares method to obtain a rigid transformation matrix T containing translation, rotation, and scaling parameters. This matrix can be decomposed into a translation vector (t x , t y ), a rotation angle θ, and a scaling factor s. For example, the transformation parameters between two adjacent frames are: translation vector (2.5, 1.8) pixels, rotation angle 1.2 degrees, and scaling factor 1.03.

[0057] Perform an affine transformation on the facial feature point coordinates to obtain the reference plane facial coordinates. Select a standard frontal facial pose as the reference plane, and map the facial feature points of each frame to this reference plane through an affine transformation to eliminate the influence caused by pose changes. The set of transformed feature points is denoted as P' = {p'1, p'2,..., p' 68}. Calculate the affine transformation matrix using the three points of the outer corners of the eyes and the tip of the nose as corresponding points, and map all feature points to the reference plane.

[0058] Calculate the feature distance matrix of the reference plane facial coordinates. On the reference plane, calculate the Euclidean distance between each pair of feature points to form a distance matrix M with a size of 68×68. Each element m ij in the matrix represents the distance between feature points p' i and p' j . By comparing the changes in the corresponding distances between different frames, the degree of facial deformation can be quantified. For example, the distance between the corners of the mouth changes from 52 pixels in frame t to 58 pixels in frame t+1, indicating that the mouth has undergone an opening deformation.

[0059] Extract the deformation parameters of the facial region. Extract the deformation parameters based on the changes in the feature distance matrix, mainly focusing on regions that are prone to deformation such as the eyes, mouth, and eyebrows. Calculate the rate of change of the distances between these regions in consecutive frames to construct a deformation parameter vector F. For example, for the mouth region, parameters such as the rate of change of the distance between the upper and lower lips and the rate of change of the distance between the corners of the mouth can be extracted to form a deformation parameter vector with a dimension of 15.

[0060] Calculate the translation component, rotation component, and scaling component in the facial pose motion model respectively. Decompose the transformation matrix T obtained previously to get the translation vector (t x , t y ), rotation angle θ, and scaling factor s. To better represent the motion trend, calculate the average value and standard deviation of 5 consecutive frames to obtain the translation component vector (t x _mean, t x _std, t y _mean, t y _std), rotation component vector (θ_mean, θ_std), and scaling component vector (s_mean, s_std).

[0061] Combine the components of the facial pose motion model with the deformation parameters of the facial region to obtain the facial motion feature vector. Concatenate the translation component vector, rotation component vector, scaling component vector, and deformation parameter vector F to form a complete facial motion feature vector V with a dimension of 23. This feature vector comprehensively expresses the rigid motion and non-rigid deformation information of the face.

[0062] Perform principal component decomposition on the face motion feature vectors to obtain the feature projection matrix. Collect face motion feature vectors from a large number of video sequences to construct a sample set. Apply principal component analysis to the sample set to extract the main directions of variation and obtain the feature projection matrix W. In practical applications, select the first 10 principal components to retain approximately 95% of the information content.

[0063] Project the face motion feature vectors onto the feature projection matrix to obtain the dimensionality-reduced feature representation. For each face motion feature vector V, perform dimensionality reduction through the projection matrix W to obtain a 10-dimensional dimensionality-reduced feature representation V'. The features after dimensionality reduction can effectively express the main patterns of face motion and deformation.

[0064] Construct a stability evaluation space based on the dimensionality-reduced feature representation. In the dimensionality-reduced feature space, define a stability metric function to calculate the stability score of each feature point in this space. The stability score takes into account factors such as the average magnitude of the feature point displacement, direction consistency, and periodicity. For example, in a 15-frame video segment, the stability score of the feature points in the corner of the eye region is 0.85 (full score is 1), 0.92 for the tip of the nose region, and 0.73 for the corner of the mouth region.

[0065] Extract the locally optimal stable intervals in the stability evaluation space. Use the sliding window method to scan the entire video sequence, with the window size set to 20 frames and the step size to 5 frames. Calculate the comprehensive stability score for each window and select the window with the highest local score as the optimal stable interval. For example, in a 100-frame video, two locally optimal stable intervals, frames 25 - 45 and frames 60 - 80, will be identified, with stability scores of 0.88 and 0.91 respectively.

[0066] Output the statistical features of the locally optimal stable intervals as the prior information of the stability of the face region. Extract statistical features such as the mean value of the feature point positions, standard deviation, and curvature of the motion trajectory within these intervals to form the prior information of stability. The output prior information includes the frame indices of the stable intervals, the average positions of each feature point, the confidence score, and the recommended best frames suitable for face recognition. These prior information can be used to improve the algorithm performance in subsequent tasks such as face recognition, 3D reconstruction, and expression analysis.

[0067] Traditional face stability assessment techniques mainly rely on image quality evaluation methods, such as static indicators like sharpness, contrast, and brightness, or simple inter-frame difference calculations. These methods usually only consider information at the image level and fail to fully utilize the motion information of facial feature points for refined analysis. For example, typical quality-based methods use the Laplacian operator to calculate image sharpness or use histogram analysis to evaluate image brightness distribution, and then select the frame with the highest quality as the best frame. Inter-frame difference-based methods usually calculate the sum of pixel differences or the Structural Similarity Index (SSIM) between adjacent frames and select the interval with the smallest change as the stable interval. These methods have limited effectiveness in face stability assessment in complex scenarios. Especially when there are facial expression changes or slight pose changes, it is difficult to accurately evaluate the true stability of the face. The method of this embodiment elevates face stability assessment from the image level to the feature point motion model level. By constructing a face pose motion model and analyzing non-rigid deformation parameters, it can more precisely describe the motion characteristics of the face. Introducing a reference plane transformation mechanism can effectively separate rigid motion and non-rigid deformation, enabling accurate assessment of the stability of facial expressions even in the case of face pose changes. Using principal component analysis for feature dimensionality reduction to construct a dedicated stability assessment space can reduce redundant information and improve computational efficiency. Designing a sliding window mechanism can discover multiple local optimal stable intervals in the video sequence, not limited to the global optimal solution, which better meets the actual application requirements.

[0068] Figure 3 It is a schematic diagram for comparing the performance of face stability assessment methods. The method proposed in the present invention has achieved significant advantages in all test scenarios. On the standard test set, the new method reaches an accuracy rate of 87.0%, which is 20.8 and 15.0 percentage points higher than the traditional method and the inter-frame difference method respectively. Especially in complex scenarios, the advantages of the new method are more obvious: it reaches an accuracy rate of 85.0% in the pose change scenario, 83.0% in the expression change scenario, and 81.0% in the illumination change scenario.

[0069] The traditional quality assessment method performs the worst in the expression change scenario (only 56.0%), indicating that it is difficult to cope with the challenges brought by non-rigid deformation. Although the inter-frame difference method has slightly better overall performance than the traditional method, it also performs poorly in the expression change scenario (60.0%). In contrast, the method proposed in the present invention based on the feature point motion model and dimensionality-reduced feature representation can more effectively evaluate face stability, especially suitable for identifying stable intervals in complex change scenarios, providing reliable prior information for subsequent face recognition and analysis tasks.

[0070] In an alternative embodiment, multi-scale features of a face image are extracted based on a pyramid decomposition structure, and a structural fingerprint code and a texture fingerprint code are generated through frequency-domain transformation compression, and a dual-index feature is formed by fusion in a feature manifold space, including: A pyramid decomposition structure is constructed for a standard face image, and a histogram of oriented gradients is extracted at each scale level of the pyramid decomposition structure, and a multi-scale structural feature matrix is obtained by cross-comparing the histograms of oriented gradients of adjacent scale levels; The gray-level co-occurrence features are calculated at each scale level of the pyramid decomposition structure, the gray-level co-occurrence features are accumulated in different directions to form a direction accumulation curve, and a multi-scale texture feature matrix is generated based on the distribution law of the direction accumulation curve; The discrete cosine transform is performed on the multi-scale structural feature matrix to obtain a frequency-domain coefficient matrix, the low-frequency components in the frequency-domain coefficient matrix are selected to construct a structural feature descriptor, and a structural fingerprint code is generated based on the energy distribution of the structural feature descriptor; The wavelet transform is performed on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, the high-frequency components in the wavelet coefficient matrix are extracted to construct a texture response sequence, and a texture fingerprint code is generated based on the energy distribution of the texture response sequence; The structural fingerprint code and the texture fingerprint code are mapped to a local feature manifold space, a cross-modal correlation matrix is constructed in the local feature manifold space, and adaptive fusion is obtained based on feature alignment loss optimization to obtain a dual-index feature.

[0071] In a specific embodiment, a standard face image is obtained and a pyramid decomposition structure is constructed for it. Specifically, the Gaussian pyramid decomposition method is used to successively downsample the original image to obtain multiple image levels with different resolutions. For a face image of 1024×1024 pixels, five scale levels such as 512×512, 256×256, 128×128, and 64×64 can be obtained through continuous downsampling. At each scale level, the Sobel operator is applied to calculate the gradient magnitude and direction of pixel points, and the histogram of oriented gradients is extracted. The edge direction interval is set from 0° to 180°, evenly divided into 12 bins, and each bin has a width of 15°. The histogram of oriented gradients at each scale level is normalized to ensure that the numerical range is between 0 and 1. The histograms of oriented gradients of adjacent scale levels are cross-compared, and the difference values of each direction bin between each pair of adjacent levels are calculated to form a difference vector. These difference vectors are combined into a multi-scale structural feature matrix with a size of (number of levels - 1) × number of direction bins, that is, 4×12.

[0072] In the texture feature extraction stage, the gray-level co-occurrence matrix is calculated for each scale level of the pyramid decomposition structure, considering four directions (0°, 45°, 90°, 135°), and the distance parameter is set to 1 pixel. Statistical features such as contrast, energy, entropy, and correlation are extracted from the gray-level co-occurrence matrix. These statistical features are accumulated in different directions to generate the direction accumulation curve. For example, for the first scale level, the accumulated value of energy in the 0° direction is 0.82, 0.75 in the 45° direction, 0.79 in the 90° direction, and 0.76 in the 135° direction. Based on the distribution law of the direction accumulation curve, the accumulated values of each statistical feature in different directions are sorted and combined to form a multi-scale texture feature matrix with the size of the number of levels × the number of features × the number of directions, that is, 5×4×4.

[0073] The discrete cosine transform is performed on the multi-scale structural feature matrix to obtain the frequency-domain coefficient matrix. The low-frequency components of the upper left 3×3 are selected from the frequency-domain coefficient matrix to construct the structural feature descriptor, which reflects the main energy distribution of the face structure. The energy values of these low-frequency components are calculated, and binary processing is performed on the energy values by setting a threshold of 0.5 to generate the structural fingerprint code. Specifically, when the energy value is greater than the threshold, it is assigned 1, otherwise 0. Taking the first scale level as an example, the extracted energy values of the low-frequency components are [0.87, 0.62, 0.43, 0.55, 0.32, 0.76, 0.41, 0.35, 0.29], and the corresponding structural fingerprint code is [1, 1, 0, 1, 0, 1, 0, 0, 0].

[0074] The wavelet transform is performed on the multi-scale texture feature matrix, and three-layer decomposition is carried out using the Daubechies wavelet basis function to obtain the wavelet coefficient matrix. The coefficients of the high-frequency sub-bands (LH, HL, HH) are extracted from the wavelet coefficient matrix to construct the texture response sequence. For each scale level, three high-frequency sub-bands can be obtained, and 5 statistics (mean, standard deviation, energy, entropy, kurtosis) are extracted from each sub-band to form a 15-dimensional feature vector. Based on the energy distribution of the texture response sequence, a dynamic threshold is set for binary processing to generate the texture fingerprint code. Taking the second scale level as an example, the energy distribution of the high-frequency sub-band LH is [0.23, 0.45, 0.67, 0.32, 0.51], and the binary code [0, 1, 1, 0, 1] is obtained after applying the threshold 0.4.

[0075] Map the structural fingerprint code and the texture fingerprint code to the local feature manifold space. Adopt the locally linear embedding algorithm, set the number of neighbor points to 10, and the embedding dimension to 20. Construct a cross-modal correlation matrix in the local feature manifold space, where the matrix elements represent the similarity between the structural features and the texture features. The similarity is calculated using the cosine distance metric, with a value range from 0 to 1. For two feature vectors of different modalities, if the cosine similarity is greater than 0.75, they are considered to have a high correlation. Optimize based on the feature alignment loss. The loss function combines the reconstruction error and the regularization term, and iteratively optimizes the parameters through the gradient descent algorithm. Set the learning rate to 0.01 and the number of iterations to 500. By minimizing the alignment loss, achieve the adaptive fusion of the structural features and the texture features to obtain the dual-index feature with a dimension of 32.

[0076] In this embodiment, the edge direction histogram and the gray-level co-occurrence features are extracted at different scale levels through the pyramid decomposition structure, which can capture both the detail and the global contour information simultaneously, and greatly improve the sensitivity and robustness to various scale features such as face deformation and expression changes; the structural fingerprint code reflects the overall facial shape and skeletal information based on the low-frequency DCT components, and the texture fingerprint code depicts the skin texture and details based on the high-frequency wavelet coefficients. The two complement each other, making the extracted features more discriminative and effectively improving the recognition accuracy; both the direction accumulation of the gray-level co-occurrence matrix and the extraction of the edge direction histogram have statistical constraints on the local gray level and edge layout, and can maintain stable feature responses under uneven illumination, weak texture or mild noise conditions, thereby enhancing the anti-interference performance of the system; map the structural and texture fingerprint codes to the local feature manifold space, and adaptively balance the importance of the two types of features through constructing a cross-modal correlation matrix and optimizing the feature alignment loss, and finally obtain a more discriminative and generalizable dual-index feature.

[0077] In an alternative embodiment, mapping the structural fingerprint code and the texture fingerprint code to the local feature manifold space, constructing a cross-modal correlation matrix in the local feature manifold space, and adaptively fusing based on the feature alignment loss optimization to obtain the dual-index feature includes: Calculate the feature distance distributions of the structural fingerprint code and the texture fingerprint code in the local neighborhood respectively to obtain the structural similarity matrix and the texture similarity matrix, and construct the feature distribution representation; Perform cross-correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain the cross-modal correlation matrix, calculate the feature weight coefficients, and perform weighted combination of the feature weight coefficients and the feature distribution representation to obtain the manifold consistency matrix; Construct a non-linear mapping model including an encoding layer and a decoding layer, input the structural fingerprint code and the texture fingerprint code, calculate the alignment loss between the mapped features based on the manifold consistency matrix, and construct an alignment loss function in combination with the reconstruction loss; Based on the alignment loss function, loss iteration optimization is carried out to obtain an optimized non-linear mapping function. The optimized non-linear mapping function is used to remap the structure fingerprint code and the texture fingerprint code to obtain remapped features. Calculate the similarity between the remapped features to obtain a feature similarity matrix, calculate the feature fusion weights in combination with the feature distribution representation, and perform weighted combination on the remapped features according to the feature fusion weights to obtain dual-index features.

[0078] In a specific implementation, the feature distance distributions of the structure fingerprint code and the texture fingerprint code in the local neighborhood are calculated respectively to obtain a structure similarity matrix and a texture similarity matrix, and a feature distribution representation is constructed. The structure fingerprint code is extracted from the key structure information of the face image, including the geometric relationships of organs such as eyes, nose, and mouth. The histogram of oriented gradients feature extraction method is used to obtain a 128-dimensional structure fingerprint code. The texture fingerprint code is extracted from the details of the face skin texture, and the local binary pattern algorithm is used to obtain a 256-dimensional texture feature vector. For each face image in the student ID photo, the local neighborhood range is defined as the 5×5 pixel area around the current feature point. Within this neighborhood, the Euclidean distance between the structure fingerprint codes is calculated to construct a 128×128 structure similarity matrix. Similarly, the Euclidean distance of the texture fingerprint codes is calculated to obtain a 256×256 texture similarity matrix. For example, for the student photo with ID S20210523, the 15th and 16th dimensional feature values of the extracted structure fingerprint code are 0.82 and 0.76 respectively, and their Euclidean distance is 0.06, which is stored in the corresponding position of the structure similarity matrix. These two matrices together constitute the feature distribution representation, which describes the distribution relationship of face features in different dimensions.

[0079] Perform cross - correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross - modal correlation matrix, calculate the feature weight coefficients, and perform weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix. The cross - correlation analysis is achieved by calculating the Pearson correlation coefficient between the features of each dimension of the structural fingerprint code and the texture fingerprint code. For the student photo with ID S20210523, the correlation coefficient between the 20th - dimension feature of the structural fingerprint code and the 45th - dimension feature of the texture fingerprint code is 0.73, indicating that the features of these two dimensions have a strong correlation. By calculating the correlation coefficients between all dimensions, a cross - modal correlation matrix of size 128×256 is constructed. Based on this matrix, the importance of each feature dimension is calculated to generate the feature weight coefficients. The weight calculation uses the centrality algorithm, and weight coefficients are generated for the 128 dimensions of the structural fingerprint code and the 256 dimensions of the texture fingerprint code respectively. For example, the weight of the 20th dimension of the structural fingerprint code is 0.85, and the weight of the 45th dimension of the texture fingerprint code is 0.78. These weight coefficients are weighted - combined with the previously obtained feature distribution representations (structural similarity matrix and texture similarity matrix) to obtain the manifold consistency matrix. This matrix describes the importance of different feature dimensions in maintaining the local geometric structure.

[0080] Construct a non - linear mapping model including an encoding layer and a decoding layer. Input the structural fingerprint code and the texture fingerprint code, calculate the alignment loss between the mapped features based on the manifold consistency matrix, and construct an alignment loss function in combination with the reconstruction loss. The non - linear mapping model adopts a multi - layer perceptron structure, including an encoding part and a decoding part. The encoding part consists of three fully - connected layers. For the structural fingerprint code, the input layer has 128 neurons, and the hidden layers have 64 and 32 neurons respectively; for the texture fingerprint code, the input layer has 256 neurons, and the hidden layers have 128 and 32 neurons respectively. The encodings of both types of fingerprint codes are finally mapped to a feature space of the same dimension (32 - dimensional). The structure of the decoding part is symmetric to the encoding part and is used to reconstruct the mapped features back to the original feature space. For the photo of student S20210523, the structural fingerprint code is encoded to obtain 32 - dimensional mapped features, and the values of its 1st and 2nd dimensions are 0.42 and 0.39 respectively. Similarly, the 1st and 2nd dimensions of the texture fingerprint code after mapping are 0.38 and 0.41 respectively. Based on the manifold consistency matrix, calculate the alignment loss between these two mapped features, mainly examining the degree of preservation of the local geometric structure before and after mapping. At the same time, calculate the reconstruction loss, that is, the difference between the encoded - decoded features and the original features. The alignment loss and the reconstruction loss are weighted - combined to obtain the final alignment loss function, where the weight of the alignment loss is 0.7 and the weight of the reconstruction loss is 0.3.

[0081] Based on the alignment loss function, loss iterative optimization is carried out to obtain an optimized non-linear mapping function. The optimized non-linear mapping function is used to remap the structural fingerprint code and the texture fingerprint code to obtain remapped features. The optimization process uses the stochastic gradient descent algorithm, with the initial learning rate set to 0.01 and halved every 5000 iterations. To prevent overfitting, an L2 regularization term is introduced, and the regularization coefficient is 0.0001. Training is performed on the student ID photo dataset for 1000 batches, with each batch containing 64 photos. After training is completed, an optimized non-linear mapping function is obtained. Taking the student photo with ID S20210523 as an example, the alignment loss before optimization is 0.385, and after 50000 iterations of optimization, the loss drops to 0.092, indicating a significant improvement in the mapping effect. Using the optimized non-linear mapping function, the structural fingerprint code and the texture fingerprint code are remapped. After remapping, the structural fingerprint code generates 32-dimensional remapped features, and the texture fingerprint code also generates 32-dimensional remapped features. For example, the first and second dimensions of the remapped structural fingerprint code become 0.44 and 0.37 respectively, and the corresponding values of the remapped texture fingerprint code are 0.43 and 0.36, indicating that the two features are closer in the mapping space.

[0082] The similarity between the remapped features is calculated to obtain a feature similarity matrix. Combining the feature distribution representation, the feature fusion weights are calculated, and the remapped features are weighted and combined according to the feature fusion weights to obtain dual-index features. The cosine similarity between the remapped features of the structural fingerprint code and the texture fingerprint code is calculated to generate a 32×32 feature similarity matrix. For student S20210523, the cosine similarity of the first dimension of the remapped structural fingerprint code and the texture fingerprint code is 0.95, indicating that these two dimensions are highly correlated in the new feature space. Based on the feature similarity matrix and the previously constructed feature distribution representation, an attention mechanism is used to calculate the feature fusion weights. Specifically, for the features of each dimension, corresponding weights are assigned according to their local geometric structure preservation degree in the feature distribution representation and their importance in the feature similarity matrix. The overall weight of the remapped features of the structural fingerprint code is 0.52, and the overall weight of the remapped features of the texture fingerprint code is 0.48, and there are also sub-weights for each dimension. For example, the weight of the first dimension of the remapped features of the structural fingerprint code is 0.058, and the weight of the first dimension of the remapped features of the texture fingerprint code is 0.053. According to these weights, the two remapped features are weighted and combined to obtain 32-dimensional dual-index features. This feature combines structural and texture information and more comprehensively represents face features, which can be used for the retrieval and identification of student ID photos. For student S20210523, the value of the first dimension of the generated dual-index feature is 0.435, which is stored in the distributed database of the student status management system for subsequent rapid retrieval and identity verification.

[0083] Table 1: Comparison table of recognition accuracies of different feature representation methods:

[0084] This solution shows significant advantages under all test conditions. Under standard lighting conditions, this solution achieves an identification accuracy rate of 96.8%, which is approximately 8 percentage points higher than traditional single-feature representation methods (structural feature 87.4% and texture feature 89.2%), and 5.3 and 4.5 percentage points higher than simple feature splicing and fixed-weight fusion methods respectively.

[0085] Particularly prominent is that this solution demonstrates excellent robustness under complex conditions. Under the severe condition of a 30° pose deflection, this solution still maintains an accuracy rate of 78.9%, while the accuracy rates of single-feature representation methods are only 62.4% and 59.3%. In the scenario of facial expression changes, the accuracy rate of this solution reaches 92.5%, far higher than other methods.

[0086] Traditional face feature extraction and student ID photo management technologies mainly adopt single-feature representation methods, such as face recognition based on structural features or texture features, and cannot make full use of multi-modal face information. Existing feature fusion methods usually use simple feature splicing or average fusion, without considering the correlation and complementarity between different features, resulting in poor fusion effects. For example, some systems directly reduce the dimension using principal component analysis and then splice different features, or use fixed weights (such as 0.5:0.5) for feature weighting, ignoring the complex relationships between features. In addition, traditional methods have poor robustness in dealing with non-linear changes (such as lighting, pose, and facial expression changes), and the changing acquisition environment of student ID photos leads to a decrease in the recognition accuracy rate.

[0087] The method of this embodiment proposes a dual feature representation of structural fingerprint codes and texture fingerprint codes to comprehensively capture the structural and texture information of faces; introduces a manifold consistency matrix to describe the local geometric structure of different feature spaces and ensure the consistency of local relationships during the feature mapping process; designs a non-linear mapping model based on alignment loss to align different modality features in a common feature space and enhance the complementarity of features; adopts an adaptive feature fusion weight calculation method to dynamically adjust the fusion weights according to feature similarity and distribution characteristics instead of simple averaging or splicing; implements a distributed storage and retrieval mechanism for student status photos, improving the concurrent processing ability and retrieval efficiency of the system; to solve problems faced in student status photo management such as large differences in photo quality, variable acquisition environments, and insufficient recognition accuracy, especially to improve recognition efficiency and accuracy in large-scale student data management scenarios. Through experimental verification, the method of this embodiment is tested on a student status photo dataset containing 50,000 students. Compared with traditional single-feature methods and simple feature splicing methods, the recognition accuracy is improved, and at the same time, the feature dimension is reduced, significantly reducing storage and computational overhead. Under non-ideal conditions (such as lighting changes and pose deflections), the method of this embodiment is significantly more robust than the comparative method, and the decrease in recognition accuracy is reduced. In addition, after adopting distributed storage, the concurrent processing ability of the system is improved, and the retrieval speed of student status photos is increased, enabling a student status management system that supports tens of thousands of people to be online simultaneously. The method of this embodiment is not only applicable to student status photo management but can also be extended to various face recognition application scenarios such as security monitoring and attendance management.

[0088] The intelligent acquisition and distributed storage management system for student status photos based on face recognition according to an embodiment of the present invention includes: A first unit for obtaining a face image video stream through a camera device and preprocessing it to obtain a preprocessed image sequence; A second unit for performing dynamic focus scanning on the preprocessed image sequence, calculating the local texture complexity and neighborhood gray uniformity of the scanning position to obtain the region of interest, performing adaptive region expansion with the highest point of the region of interest as the center to obtain face positioning parameters, and extracting the target face image; A third unit for extracting facial key points based on the target face image according to preset points, and calculating state evaluation parameters including facial pose angle values, eye opening degree values, and expression parameter values based on the facial key points; A fourth unit for calculating the phase spectrum change feature of consecutive target face images, determining the state stability score in combination with prior stability information, and generating a shooting trigger signal based on an adaptive threshold; A fifth unit for responding to the shooting trigger signal, collecting face photos and performing enhancement processing to obtain standard face images; The sixth unit is used to extract multi-scale features of a face image based on a pyramid decomposition structure, compress and generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation, and fuse them in a feature manifold space to form a dual-index feature; The seventh unit is used to store the standard face image and the dual-index feature in a distributed storage system and generate an electronic voucher containing the storage location.

[0089] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0090] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0091] The present invention can be a method, a device, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent acquisition and distributed storage management method for school roll photos based on face recognition, characterized in that Including: Obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence; Perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood gray level uniformity of the scanning position to obtain the regional interest degree, perform adaptive regional expansion with the highest point of the regional interest degree as the center to obtain face positioning parameters, and extract the target face image; Based on the target face image, extract facial key points according to preset points, and calculate state evaluation parameters including facial pose angle values, eye opening degrees, and expression parameter values based on the facial key points; Calculate the phase spectrum change characteristics of consecutive target face images, combine the stability prior information to determine the state stability score, and generate a shooting trigger signal based on an adaptive threshold; In response to the shooting trigger signal, collect a face photo and perform enhancement processing to obtain a standard face image; Extract multi-scale features of the face image based on the pyramid decomposition structure, generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation compression, and fuse them in the feature manifold space to form a dual-index feature; Store the standard face image and the dual-index feature in a distributed storage system, and generate an electronic voucher containing the storage location.

2. The method according to claim 1, wherein Performing dynamic focus scanning on the preprocessed image sequence and calculating the local texture complexity and neighborhood gray level uniformity of the scanning position to obtain the regional interest degree includes: Establish a dynamic angle step size based on the image center point coordinates of the preprocessed image sequence, multiply the dynamic angle step size by a dynamic expansion factor to generate a dynamic radius step size, and calculate the initial scanning coordinates according to the dynamic angle step size and the dynamic radius step size; Calculate the local response intensity value based on the initial scanning coordinates, generate a focus adjustment coefficient by taking the ratio of the local response intensity value to a preset reference threshold, and adjust the sampling position of the initial scanning coordinates according to the focus adjustment coefficient to generate a dynamic focus scanning position; Select a first pixel window based on the dynamic focus scanning position, generate the local texture complexity by calculating the cumulative difference between the pixel points in the first pixel window and the weighted average value of the window, select a second pixel window based on the dynamic focus scanning position, and generate the neighborhood gray level uniformity by calculating the ratio of the pixel dispersion degree of the second pixel window to the overall dispersion degree of the image; Multiply the local texture complexity and the neighborhood gray level uniformity by the corresponding preset weights respectively to obtain a first weighted feature and a second weighted feature, determine the basic region feature value, calculate the boundary attenuation coefficient according to the distance from the dynamic focus scanning position to the image boundary, and generate the regional interest degree by multiplying the basic region feature value by the boundary attenuation coefficient.

3. The method according to claim 1, wherein Performing adaptive regional expansion with the highest point of the regional interest degree as the center to obtain face positioning parameters and extracting the target face image includes: Determine the region growth starting point at the maximum value position of the regional interest degree, construct a polar coordinate grid with the region growth starting point as the center, calculate the interest degree gradient value on each direction axis of the polar coordinate grid, generate the axial expansion distance by multiplying the interest degree gradient value by the length of the direction axis, determine the boundary point coordinates according to the axial expansion distance, and perform low-pass filtering on the sequence composed of the boundary point coordinates to obtain a smooth region contour; Perform elliptical fitting on the smooth region contour, calculate the region center coordinates, major axis length, minor axis length, and orientation angle, and combine them to generate a set of positioning parameters. Construct a coordinate transformation matrix based on the positioning parameter set, apply the coordinate transformation matrix to the preprocessed image sequence for geometric mapping, and perform size normalization processing to extract the target face image.

4. The method according to claim 1, wherein Calculate the phase spectrum change features for consecutive target face images, combine the stability prior information to determine the state stability score, and generate a shooting trigger signal based on an adaptive threshold, including: Perform two-dimensional Fourier transform on consecutive target face images to obtain a frequency domain representation, extract the phase components from the frequency domain representation to generate a phase spectrum sequence, and stack the phase spectrum sequence in the time dimension to form a phase spectrum tensor; Calculate the phase spectrum difference between adjacent time frames in the phase spectrum tensor to obtain a phase spectrum difference sequence, construct a frequency weight function based on a Gaussian kernel, and perform weighted combination of the frequency weight function and the phase spectrum difference sequence to obtain the phase spectrum change features; Obtain the stability prior information of the face region, map the stability prior information to the frequency space to obtain a state evaluation parameter, and generate a state stability score according to the combined distribution law of the phase spectrum change features and the state evaluation parameter; Calculate the mean of the state stability scores within a historical time window, determine the dynamic adjustment coefficient as the ratio of the mean to a preset stability threshold, and adjust the preset stability threshold according to the dynamic adjustment coefficient to obtain an adaptive decision threshold; generate a shooting trigger signal when the state stability scores are all greater than the adaptive decision threshold within a consecutive preset frame number threshold.

5. The method according to claim 4, characterized in that, Obtaining the stability prior information of the face region includes: Obtain the coordinates of the face feature points in the target face image, calculate the displacement vector of the face feature point coordinates between adjacent image frames, and construct a face pose motion model based on the displacement vector; Perform affine transformation on the face feature point coordinates to obtain the reference plane face coordinates, calculate the feature distance matrix of the reference plane face coordinates, and extract the face region deformation parameters; Calculate the translation component, rotation component, and scaling component in the face pose motion model respectively, and combine them with the face region deformation parameters to obtain a face motion feature vector; Perform principal component analysis on the face motion feature vector to obtain a feature projection matrix, project the face motion feature vector onto the feature projection matrix to obtain a dimensionality-reduced feature representation, and construct a stability evaluation space based on the dimensionality-reduced feature representation; Extract the local optimal stable interval in the stability evaluation space, and output the statistical features of the local optimal stable interval as the stability prior information of the face region.

6. The method according to claim 1, characterized in that Extract multi-scale features of the face image based on a pyramid decomposition structure, generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation compression, and fuse them in the feature manifold space to form a dual-index feature, including: Construct a pyramid decomposition structure for a standard face image, extract the histogram of oriented gradients at each scale level of the pyramid decomposition structure, and perform cross-comparison of the histograms of oriented gradients at adjacent scale levels to obtain a multi-scale structure feature matrix; Calculate the gray-level co-occurrence features at each scale level of the pyramid decomposition structure, accumulate the gray-level co-occurrence features in different directions to form a direction accumulation curve, and generate a multi-scale texture feature matrix based on the distribution law of the direction accumulation curve; Perform a discrete cosine transform on the multi-scale structural feature matrix to obtain a frequency-domain coefficient matrix, select the low-frequency components in the frequency-domain coefficient matrix to construct a structural feature descriptor, and generate a structural fingerprint code based on the energy distribution of the structural feature descriptor; Perform a wavelet transform on the multi-scale texture feature matrix to obtain a wavelet coefficient matrix, extract the high-frequency components in the wavelet coefficient matrix to construct a texture response sequence, and generate a texture fingerprint code based on the energy distribution of the texture response sequence; Map the structural fingerprint code and the texture fingerprint code to the local feature manifold space, construct a cross-modal correlation matrix in the local feature manifold space, and perform adaptive fusion based on the optimization of the feature alignment loss to obtain a dual-index feature.

7. The method according to claim 6, wherein Mapping the structural fingerprint code and the texture fingerprint code to the local feature manifold space, constructing a cross-modal correlation matrix in the local feature manifold space, and performing adaptive fusion based on the optimization of the feature alignment loss to obtain a dual-index feature includes: Calculate the feature distance distributions of the structural fingerprint code and the texture fingerprint code in the local neighborhood respectively to obtain a structural similarity matrix and a texture similarity matrix, and construct a feature distribution representation; Perform a cross-correlation analysis on the local features of the structural fingerprint code and the texture fingerprint code to obtain a cross-modal correlation matrix, calculate the feature weight coefficients, and perform a weighted combination of the feature weight coefficients and the feature distribution representation to obtain a manifold consistency matrix; Construct a non-linear mapping model including an encoding layer and a decoding layer, input the structural fingerprint code and the texture fingerprint code, calculate the alignment loss between the mapping features based on the manifold consistency matrix, and construct an alignment loss function in combination with the reconstruction loss; Based on the alignment loss function, perform loss iteration optimization to obtain an optimized non-linear mapping function, and use the optimized non-linear mapping function to remap the structural fingerprint code and the texture fingerprint code to obtain remapped features; Calculate the similarity between the remapped features to obtain a feature similarity matrix, calculate the feature fusion weights in combination with the feature distribution representation, and perform a weighted combination of the remapped features according to the feature fusion weights to obtain a dual-index feature.

8. A smart collection and distributed storage management system for student status photos based on face recognition, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to obtain a face image video stream through a camera device and preprocess it to obtain a preprocessed image sequence; The second unit is used to perform dynamic focus scanning on the preprocessed image sequence, calculate the local texture complexity and neighborhood gray-level uniformity of the scanning position to obtain the region of interest, perform adaptive region expansion with the highest point of the region of interest as the center to obtain face positioning parameters, and extract the target face image; The third unit is used to extract facial key points based on the target face image according to preset points, and calculate state evaluation parameters including facial pose angle values, eye opening degrees, and expression parameter values based on the facial key points; The fourth unit is used to calculate the phase spectrum change features of consecutive target face images, determine the state stability score in combination with the stability prior information, and judge and generate a shooting trigger signal based on an adaptive threshold; The fifth unit is configured to collect a face photo in response to a shooting trigger signal and perform enhancement processing to obtain a standard face image; The sixth unit is configured to extract multi-scale features of a face image based on a pyramid decomposition structure, generate a structure fingerprint code and a texture fingerprint code through frequency domain transformation compression, and fuse them in a feature manifold space to form a dual-index feature; The seventh unit is configured to store the standard face image and the dual-index feature in a distributed storage system and generate an electronic voucher including the storage location.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for detecting classroom attention of student

    CN104517102A

  • Face expression identification method based on video sequences

    CN105139004A

  • Human face tracking optimization method for camera and intelligent health monitoring system based on videos

    CN105868574A

  • Rapid face detection identification method based on deep learning

    CN108564049A

  • Face recognition method and system

    CN109753904A

Cited By

  • Face depth detection method and system based on multi-mode double shooting

    CN121096003A