Intelligent conference participant automatic marking method based on image recognition

Through technologies such as multi-angle image acquisition and lighting compensation, anti-occlusion feature vectors are extracted and matched with the identity database, which solves the identity identification problem in the case of occlusion of participants, and ensures data security through encrypted tagged data packets, achieving efficient identity labeling and permission management.

CN120198952AActive Publication Date: 2025-06-24广东公信智能会议股份有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510689531.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The prior art cannot achieve accurate and stable identity identification and permission control when faced with complex scenarios such as participants wearing masks, glasses or occlusions. In addition, the generation and transmission of tagged data in the conference system does not use an encryption mechanism, which poses a security risk of data leakage.

Method used

The multi-angle image acquisition device acquires the face images of participants in real time, performs lighting compensation, posture correction and occlusion area segmentation processing, extracts anti-occlusion feature vectors, and matches them with the pre-constructed participant identity database to generate encrypted tagged data packets to realize identity identification and permission management.

Benefits of technology

In the presence of occlusion of participants, stable facial recognition and identity marking are achieved, which improves the automation level of identity marking and permission management efficiency at the conference site, and ensures data security through encryption mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198952A_ABST
    Figure CN120198952A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition and conference management, in particular to an intelligent conference participant automatic marking method based on image recognition, which comprises the following steps: S1, acquiring facial images of participants in real time, and generating structured participant image data; s2, extracting facial features of the participants from the structured image data of the participants, and generating an anti-shielding feature vector; s3, matching the anti-shielding feature vector with a pre-constructed participant identity database; s4, associating a preset permission rule in the conference management system, and generating composite mark information; s5, performing homomorphic encryption processing on the composite mark information; and S6, decrypting the encrypted marked data packet, and superposing the decrypted encrypted marked data packet to the real-time video stream of the participants. According to the invention, through anti-shielding feature extraction, authority automatic association and encryption dynamic labeling mechanisms, the accuracy of identity recognition of participants, the real-time performance of authority management and the security of data transmission in a complex conference scene are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image recognition and conference management, and particularly to an automatic marking method for intelligent conference participants based on image recognition. Background Art

[0002] With the popularization of multi-party remote collaborative work and large-scale offline conferences, conference organizers have put forward higher requirements for the real-time recognition of the identities of participants and permission management; traditional conference sign-in and identity verification mainly rely on manual verification or radio frequency card recognition, which not only has low efficiency and is error-prone, but also cannot achieve accurate and stable identity recognition and permission control in the face of complex scenarios such as participants wearing masks, glasses or having occlusions.

[0003] Most of the existing image recognition technologies are based on full-face feature extraction, lacking robust processing of local features and being unable to adapt to non-ideal shooting conditions such as partial facial occlusion, uneven illumination or pose deviation; in addition, the generation and transmission of marked data in the current conference system mostly do not adopt an encryption mechanism, posing a security risk of data leakage. Therefore, there is an urgent need for an automatic marking method for intelligent conference participants based on image recognition to solve the above problems. Summary of the Invention

[0004] Based on the above purpose, the present invention provides an automatic marking method for intelligent conference participants based on image recognition.

[0005] The automatic marking method for intelligent conference participants based on image recognition includes the following steps: S1: Through multi-angle image acquisition devices deployed in the conference scene, the facial images of participants are obtained in real time, and the facial images are subjected to illumination compensation, pose correction and occlusion area segmentation processing to generate structured participant image data; S2: Extract the facial features of the participants from the structured participant image data. For the scenarios of wearing masks and glasses occlusion, a local feature enhancement algorithm is adopted to extract the geometric distribution features of the eye area and the texture features of the nose bridge area to generate an anti-occlusion feature vector; S3: Match the anti-occlusion feature vector with a pre-constructed participant identity database, and the participant identity database includes encrypted biometric codes, position information and permission tags of registered participants; If the match is successful, the corresponding identity tag data is output; If the match fails, an exception marking for unregistered personnel is triggered to generate temporary identity tag data including a temporary number and an admission time; S4: According to the identity tag data or the temporary identity tag data, associate the preset permission rules in the conference management system to generate composite marking information including name, position, permission level and exception status identifier; S5: Perform homomorphic encryption processing on the composite marker information to generate an encrypted marker data packet; S6: In the interactive interface of the conference management terminal, decrypt the encrypted marker data packet and overlay it on the real-time video stream of the participants, and trigger dynamic refresh and repositioning of the marker information based on events such as the movement of the participants' positions or changes in permissions.

[0006] Optionally, the specific steps of S1 include: S11: Synchronously trigger shooting through the camera groups preset at the four corners and the center of the conference scene to cover the full-angle facial area of the participants, and generate an original multi-view image set; S12: Perform light intensity analysis on the original multi-view image set. For overexposed or underexposed areas in a single-frame image, adopt a multi-frame fusion compensation algorithm to extract the optimal light pixels in each area and generate a light-balanced image; S13: Based on the light-balanced image, detect the coordinates of facial key points, calculate the spatial vector angle between the tip of the nose point and the points of the two earlobes to determine the head deflection angle; when the deflection angle exceeds the preset threshold, use affine transformation to correct the facial image to a frontal view and generate a posture-normalized image; S14: Perform edge detection on the posture-normalized image, identify the contour of the mask, glasses or hand occlusion area, and generate a repaired and filled image for the occlusion area according to the facial skin texture features that are not occluded; S15: Bind the repaired and filled image with the associated acquisition timestamp, camera position code, light compensation parameter, and posture correction parameter, and encapsulate it into structured participant image data containing multi-dimensional metadata.

[0007] Optionally, the specific steps of S2 include: S21: Receive the structured participant image data generated in S1, locate the facial area in the image, and use a preset facial key point detection model to mark key points such as the corners of the eyes, the ends of the eyebrows, the center of the eyebrows, the tip of the nose, and the root of the nose, and delimit the boundary range of the periorbital area and the nasal bridge area; S22: Calculate the geometric relationships between key points within the delimited periorbital area, including the Euclidean distance and relative angle between the end of the eyebrow and the center of the eyebrow, and between the corners of the eyes, to construct the local geometric distribution features of the face; S23: In the nasal bridge area, use the local texture gradient analysis method to analyze the gray-scale change distribution in this area, extract the gray-scale gradient information representing skin texture differences, and extract multi-scale local texture features in a fixed window sliding manner; S24: Stitch and fuse the geometric distribution features and texture features obtained in S22 and S23 to generate a one-dimensional anti-occlusion feature vector.

[0008] Optionally, the specific steps of S21 include: S211: Based on the structured participant image data, obtain the image coordinates of the outer corner points of the eyes, the end points of the eyebrows, the midpoint of the eyebrows, the tip of the nose, and the root of the nose through a pre-trained facial key point detection model; S212: According to the detected key point coordinates, define the bounding box of the periorbital region. Taking the outer corner points of the left and right eyes as the center, the width of the box is 1.5 times the Euclidean distance between the two outer corner points, and the height is 0.8 times the vertical distance from the midpoint of the eyebrows to the tip of the nose. The center line of the box coincides with the line connecting the two outer corner points; S213: Define the bounding box of the nasal bridge region. Taking the midpoint between the root of the nose and the tip of the nose as the center of the box, the height of the box is set to 1.2 times the Euclidean distance between the root of the nose and the tip of the nose, and the width of the box is 0.6 times the horizontal distance between the two inner corner points. The vertical center line of the box coincides with the line connecting the root of the nose and the tip of the nose.

[0009] Optionally, the S22 specifically includes: S221: Using the facial image coordinates obtained in S21, with the midpoint of the eyebrows as the reference, calculate the Euclidean distances between the midpoint of the eyebrows and the end point of the left eyebrow, the midpoint of the eyebrows and the end point of the right eyebrow, and the outer corner point of the left eye and the outer corner point of the right eye respectively, to obtain the characteristic lengths of the periorbital region; S222: Calculate the angles formed by connecting the midpoint of the eyebrows with the end points of the left and right eyebrows respectively, and the angles formed by connecting the midpoint of the eyebrows with the outer corner points of the left and right eyes respectively, to obtain the characteristic angles of the periorbital region; S223: Normalize the Euclidean distances obtained in S221 and the angles obtained in S222; S224: Concatenate the normalized characteristic lengths and characteristic angles in sequence to form the local geometric distribution feature of the face.

[0010] Optionally, the S23 specifically includes: S231: Extract the grayscale image sub-region from the bounding box of the nasal bridge region delimited in S21, denoted as the region to be analyzed; S232: Taking each pixel point in the region to be analyzed as the center, calculate the gradient amplitude of the grayscale value along the horizontal and vertical directions of the image respectively, to obtain the local gradient intensity information; S233: Based on the gradient amplitude, use a square window with a preset size to slide and traverse the region to be analyzed at a fixed step length. The histogram distribution of the gradient amplitude is statistically calculated within each window to form the local texture feature histogram; S234: Change the window size, repeat S233, extract the gradient histograms respectively with different window sizes, and concatenate the histogram features under different window sizes in sequence to form the local texture feature with multi-scale information.

[0011] Optionally, the S3 specifically includes: S31: Extract the encrypted biometric codes of registered participants from the pre-built participant identity database, decrypt the encrypted biometric codes, and convert them into identity feature vectors in standard format; S32: Calculate the similarity between the anti-occlusion feature vector obtained in S2 and each identity feature vector one by one, and calculate the matching similarity values between each identity feature vector and the anti-occlusion feature vector through the cosine similarity method; S33: Sort the matching similarity values obtained in S32, select the identity feature vector with the highest matching similarity, and determine whether the similarity value is higher than the preset identity matching threshold; S34: If the highest matching similarity value in S33 exceeds the identity matching threshold, it is determined that the match is successful, and the identity label data associated with the corresponding identity feature vector is output; if the highest matching similarity value does not exceed the identity matching threshold, it is determined that the match fails, and the unregistered person exception marking program is triggered to output the temporary identity label data including the temporary number and the admission time.

[0012] Optionally, the specific steps of S4 are as follows: S41: Receive the identity label data or temporary identity label data output by S3, and retrieve the corresponding identity attribute information according to the identity number in the identity label data or the temporary number in the temporary identity label data; S42: Based on the preset permission rule table in the conference management system, use the position information in the identity attribute information retrieved in S41 as an index to automatically determine the permission level of the participant; S43: Determine whether the current participant identity data comes from the temporary identity label data. If the participant data comes from the temporary identity label data, add an exception status flag for the participant; if it comes from the registered identity label data, set the exception status flag to the default normal status; S44: Merge the name and position information obtained in S41, the permission level obtained in S42, and the exception status flag determined in S43 to form a structured data record and generate composite marking information.

[0013] Optionally, the specific steps of S5 are as follows: S51: Receive the composite marking information generated in S4, perform homomorphic encryption processing using the preset public key, encrypt each data field including name, position, permission level, and exception status flag item by item to generate corresponding ciphertexts; S52: After splicing the ciphertexts of each field generated in S51 in a predetermined order, add a predefined data packet header and a checksum tail mark, and encapsulate them into a complete encrypted marking data packet that conforms to the conference management terminal parsing protocol; S53: Send the encrypted tag data packet completed in S52 to the conference management terminal through a preset independent communication channel using a secure communication protocol. This communication channel is physically isolated from the video stream communication channel and is only used to transmit encrypted identity tag information data. S54: After receiving the encrypted tag data packet, the conference management terminal checks the headers and tails of the data packet according to the agreed parsing protocol.

[0014] Optionally, the S6 specifically includes: S61: After the conference management terminal receives the encrypted tag data packet transmitted in S5, homomorphically decrypt each ciphertext field in the data packet using the pre-stored private key to obtain the plaintext data of the composite tag information. S62: Automatically associate the decrypted composite tag information with the corresponding participant image position coordinates in the real-time video stream of the participant according to the participant identity number or temporary number. After performing image coordinate matching, superimpose and display the name, position, permission level, and exception status identifier near the corresponding participant position. S63: Continuously monitor the position information of the participant. If it is detected that the position coordinates of the participant have changed compared to the previous moment, trigger a position movement event, recalculate the display coordinates of the tag information in the video frame, and immediately update the superimposed position. S64: Continuously monitor the status information of the permission level of the participant in the conference management system. If the system detects that the permission level of the participant has changed, update the data content of the composite tag information according to the latest permission level and synchronously refresh the superimposed tag displayed on the video frame. S65: Continuously execute S63 and S64 to ensure that the participant tag information displayed in the real-time video stream always remains consistent with the actual position and real-time permission status of the participant.

[0015] Advantages of the present invention: In the present invention, through multi-angle image acquisition and preprocessing means, combined with a local feature enhancement algorithm, it is possible to stably extract the key structures and texture information of the eye and nose regions in the case of the participant wearing a mask, glasses, or hand occlusion, etc., generate a robust anti-occlusion feature vector, and improve the accuracy and adaptability of face recognition; at the same time, adopt a multi-scale texture analysis and geometric distribution feature fusion method to further enhance the expression ability of features under complex illumination and pose change conditions, ensuring reliable and stable identity recognition results in the actual conference environment.

[0016] In the present invention, by automatically associating the recognized identity information with the permission rules in the conference management system, a composite tag information including name, position, permission level, and exception status identifier is constructed, and data encapsulation and transmission are completed through homomorphic encryption and an independent communication channel, realizing the secure release and real-time update of identity data; the superimposed tag content can be dynamically repositioned and refreshed according to the movement of the attendee's position and the change of permissions, thus significantly improving the automation level of identity annotation and the efficiency of permission management at the conference site. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only for the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 Schematic diagram of the automatic marking method for intelligent conference participants in the embodiment of the present invention; Figure 2 Schematic diagram of the method for generating anti-occlusion feature vectors in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for a more specific description of the embodiments and are not intended to specifically limit the present invention.

[0020] It should be pointed out that in the specification, the mention of "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicates that the described embodiment may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. In addition, when combining an embodiment to describe a specific feature, structure, or characteristic, implementing such a feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0021] Generally, the terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. In addition, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.

[0022] As Figure 1 - Figure 2 shown, the intelligent conference participant automatic marking method based on image recognition includes the following steps: S1: Through multi-angle image acquisition devices deployed in the conference scene, the facial images of the participants are obtained in real time, and the facial images are subjected to light compensation, pose correction, and occlusion area segmentation processing to generate structured participant image data; S2: Extract the facial features of the participants from the structured participant image data. For the scenarios where the face is covered by a mask or glasses, a local feature enhancement algorithm is used to extract the geometric distribution features of the eye area and the texture features of the nose bridge area to generate an anti-occlusion feature vector; S3: Match the anti-occlusion feature vector with a pre-constructed participant identity database, which includes the encrypted biometric codes, position information, and permission labels of the registered participants; If the match is successful, the corresponding identity label data is output; If the match fails, an exception marking for unregistered personnel is triggered, and temporary identity label data including a temporary number and the entry time is generated; S4: According to the identity label data or the temporary identity label data, associate the preset permission rules in the conference management system to generate composite marking information including name, position, permission level, and exception status identifier; S5: Perform homomorphic encryption processing on the composite marking information to generate an encrypted marking data packet; S6: In the interactive interface of the conference management terminal, decrypt the encrypted marking data packet and overlay it on the real-time video stream of the participants, and trigger the dynamic refresh and repositioning of the marking information based on the events of the participant's position movement or permission change.

[0023] S1 specifically includes: S11: Through the camera groups preset at the four corners and the center of the conference scene, synchronously trigger shooting to cover the full-angle facial area of the participants, and generate an original multi-view image set; S12: Perform light intensity analysis on the original multi-view image set. For the overexposed or underexposed areas in a single-frame image, a multi-frame fusion compensation algorithm is used to extract the optimal light pixel values of each area to generate a light-balanced image; Specifically, let the gray value of the pixel point in the frame image be , and its local light confidence level be , then the pixel value calculation formula of the light-balanced image is: , where represents the pixel value of the fused image at the position , is the total number of frames; this formula can achieve the smoothing of local illumination extreme pixels and improve the overall brightness uniformity of the image; S13: Based on the illumination equalized image, detect the coordinates of facial key points, and determine the head deflection angle by calculating the spatial vector angle between the tip of the nose point and the points of both earlobes; when the deflection angle exceeds the preset threshold, use affine transformation to correct the facial image to a frontal view and generate a pose-normalized image; specifically, let the coordinates of the tip of the nose point be and the left ear point and right ear point be , then the head deflection angle is calculated by the following formula: , if (where is the set threshold), then execute the affine transformation matrix to correct the image, and the calculation formula is: ; where, is the rotation angle, is the translation amount, ensuring that the face in the corrected image faces the camera directly; S14: Perform edge detection on the pose-normalized image, identify the contours of the mask, glasses or hand occlusion areas, and generate a repaired and filled image of the occlusion area according to the facial skin texture features of the unoccluded area; S15: Bind the repaired and filled image with the associated acquisition timestamp, camera position encoding, illumination compensation parameters and pose correction parameters, and encapsulate it into structured participant image data containing multi-dimensional metadata. The above steps improve the image illumination uniformity through multi-frame fusion, supplemented by pose correction and occlusion repair based on key point calculation, which can ensure that even in low light, side face or occlusion situations, unified, complete and recognizable structured participant image data can be generated, significantly enhancing the accuracy and robustness of subsequent face recognition and identity marking.

[0024] S2 specifically includes: S21: Receive the structured participant image data generated in S1, locate the facial area in the image, and use a preset face key point detection model to mark key points such as the corners of the eyes, the ends of the eyebrows, the center of the eyebrows, the tip of the nose and the root of the nose, and delimit the boundary ranges of the periorbital area and the nasal bridge area; S22: In the delimited periorbital area, calculate the geometric relationships between key points, including the Euclidean distance and relative angle between the end of the eyebrow and the center of the eyebrow, and between the corners of the eyes, and construct the local geometric distribution features of the face for identifying the stable structure of individual features; S23: In the nasal bridge area, use the local texture gradient analysis method to analyze the gray level change distribution in this area, extract the gray level gradient information representing skin texture differences, and extract multi-scale local texture features in a fixed window sliding manner; S24: Concatenate and fuse the geometric distribution features and texture features obtained in S22 and S23 to generate a one-dimensional anti-occlusion feature vector. This feature vector has recognition robustness against occlusion areas such as glasses and masks, and retains the identity feature information of the areas on the face that are not easily occluded. Through the above steps, by extracting regional features from the periorbital and nasal bridge areas in the structured image and fusing the local geometric structure and texture difference information, the interference of the occlusion area on the overall recognition can be effectively bypassed, significantly improving the integrity and recognition accuracy of facial feature expression under occlusion conditions, and providing a stable input basis for anti-occlusion identity recognition.

[0025] Specifically, S21 includes: S211: Based on the structured participant image data, obtain the image coordinates of the outer corner points of the eyes, the end points of the eyebrows, the center point of the eyebrows, the tip point of the nose, and the root point of the nose through a pre-trained face key point detection model. S212: According to the detected key point coordinates, define the bounding box of the periorbital area. With the outer corner points of the left and right eyes as the center, the width of the box is taken as 1.5 times the Euclidean distance between the two outer corner points, and the height is taken as 0.8 times the vertical distance from the center point of the eyebrows to the tip point of the nose. The center line of the box coincides with the line connecting the two outer corner points. S213: Define the bounding box of the nasal bridge area. Take the midpoint between the root point of the nose and the tip point of the nose as the center of the box. The height of the box is set as 1.2 times the Euclidean distance between the root point of the nose and the tip point of the nose, and the width of the box is 0.6 times the horizontal distance between the two inner corner points. The vertical center line of the box coincides with the line connecting the root point of the nose and the tip point of the nose. Specifically, the width and height of the periorbital area bounding box are calculated by the following formulas respectively: ; ; The width and height of the nasal bridge area bounding box are calculated by the following formulas respectively: ; where represent the coordinates of the left and right outer corner points of the eyes respectively, represent the coordinates of the left and right inner corner points of the eyes respectively, and represent the coordinates of the center point of the eyebrows, the tip point of the nose, and the root point of the nose respectively. Through the above steps, the boundaries of the periorbital and nasal bridge areas are delimited by clear geometric formulas, which can accurately locate stable and reliable local facial areas, provide an accurate spatial basis for subsequent local feature extraction, and further improve the consistency and anti-occlusion robustness of feature extraction.

[0026] Obtaining the image coordinates of the outer corner points of the eyes, the end points of the eyebrows, the midpoint between the eyebrows, the tip of the nose, and the root of the nose through a pre-trained facial key point detection model includes the following steps: S2111: Input the structured participant image data into the pre-trained facial key point detection model. This model is based on a convolutional neural network structure and has been trained with a large number of labeled facial training sets, possessing the ability to predict facial key points with high accuracy; S2112: The pre-trained model first extracts features from the input image, outputs multi-scale facial feature maps, and then uses a fully connected regression layer to perform regression prediction on the feature maps, outputting the two-dimensional coordinate values of each key point; S2113: According to the normalized coordinate values output by the model and combined with the actual size of the input image, perform scale mapping on the normalized coordinates of the outer corner points of the eyes, the end points of the eyebrows, the midpoint between the eyebrows, the tip of the nose, and the root of the nose to obtain the precise coordinate positions of each key point in the image coordinate system; The scale mapping calculation formula for the actual image coordinates of the key points is: , where is the actual coordinate of the key point in the image, is the normalized coordinate output by the model, and are the actual width and height of the input image respectively.

[0027] S22 specifically includes: S221: Using the facial image coordinates obtained in S21, with the midpoint between the eyebrows as the reference, calculate the Euclidean distances between the midpoint between the eyebrows and the left end point of the eyebrow, the midpoint between the eyebrows and the right end point of the eyebrow, and the left outer corner point of the eye and the right outer corner point of the eye respectively to obtain the characteristic lengths of the periorbital region; specifically, let the Euclidean distance be , and its calculation formula is: , where in the formula, represents the image coordinate of the starting point with the midpoint between the eyebrows as the reference in the structured image; represents the image coordinate of the target point corresponding to the midpoint between the eyebrows in the structured image; S222: Calculate the angles formed by connecting the midpoint between the eyebrows with the left and right end points of the eyebrows and the angles formed by connecting the midpoint between the eyebrows with the left and right outer corner points of the eyes respectively to obtain the characteristic angles of the periorbital region; Let the angle be , and its calculation formula is: , where in the formula, represents the image coordinate of the midpoint between the eyebrows; represents the coordinate of the left key point corresponding to the midpoint between the eyebrows as the vertex; represents the coordinate of the right key point corresponding to the midpoint between the eyebrows as the vertex; the symbol • represents the dot product operation of two vectors; the symbol Represents the modulus of a vector; S223: Normalize the Euclidean distance obtained in S221 and the included angle obtained in S222 to eliminate the influence of different participant image size differences on feature stability; The length normalization formula is: ; The included angle normalization formula is: ; In the formula, and are the normalized length value and included angle value respectively; and represent the upper limit value and lower limit value of the preset length normalization respectively; and represent the upper limit value and lower limit value of the preset included angle normalization respectively; S224: Concatenate the normalized feature length and feature included angle in sequence to form a facial local geometric distribution feature descriptor, which provides geometric structure feature input for the subsequent anti-occlusion feature vector; Through the above steps, the stability of local geometric features against occlusion interference can be significantly enhanced, ensuring the consistency and robustness of feature expression, thereby improving the accuracy of identity recognition.

[0028] S23 specifically includes: S231: Extract a grayscale image sub-region from the nose bridge region bounding box delimited in S21, denoted as the region to be analyzed; S232: Calculate the gradient magnitude of the grayscale value along the horizontal and vertical directions of the image respectively with each pixel point in the region to be analyzed as the center, to obtain local gradient intensity information; Gradient magnitude The calculation formula is: , where, represents the grayscale change gradient value along the horizontal direction of the image, represents the grayscale change gradient value along the vertical direction of the image; and are calculated respectively through the following Sobel operators: ; ; Among them, is the grayscale value of the pixel point in the region to be analyzed; S233: Based on the gradient magnitude, use a square window of a preset size to slide and traverse the region to be analyzed at a fixed step length, and statistically analyze the histogram distribution of the gradient magnitude within each window to form a local texture feature histogram; The statistical method of the gradient magnitude histogram is: , where, is the histogram frequency when the gradient magnitude is an integer value ; is the current sliding window area; is an indicator function that takes the value 1 when the condition in the parentheses holds, and 0 otherwise; round represents rounding the gradient magnitude; S234: Change the window size, repeat S233, extract the gradient histograms with different window sizes respectively, and sequentially splice the histogram features under different window sizes to form a local texture feature descriptor with multi-scale information for constructing the subsequent anti-occlusion feature vector; The above steps can capture rich local texture details through precise calculation of the local gray gradient in the bridge of the nose area and multi-scale window analysis, effectively improve the discrimination ability and stability of the texture features for identifying occlusion areas, and provide a highly reliable local texture feature basis for the identification of the identities of participants.

[0029] S3 specifically includes: S31: Extract the encrypted biometric feature codes of registered participants from the pre-constructed participant identity database, and decrypt and convert the encrypted biometric feature codes into identity feature vectors in standard format; S32: Calculate the similarity between the anti-occlusion feature vector obtained in S2 and each identity feature vector one by one, and calculate the matching similarity values between each identity feature vector and the anti-occlusion feature vector through the cosine similarity method; S33: Sort the matching similarity values obtained in S32, select the identity feature vector with the highest matching similarity, and determine whether the similarity value is higher than the preset identity matching threshold; S34: If the highest matching similarity value in S33 exceeds the identity matching threshold, it is determined that the match is successful, and the identity label data associated with the corresponding identity feature vector is output; If the highest matching similarity value does not exceed the identity matching threshold, it is determined that the match fails, and the unregistered person exception marking program is triggered to output the temporary identity label data including the temporary number and the entry time; The above steps can effectively ensure the objectivity and accuracy of the identity recognition process through a clear feature vector similarity calculation and threshold matching process, improve the accurate recognition efficiency of the system for the identities of participants under occluded facial conditions, and avoid misjudgment and missed judgment phenomena.

[0030] Table 1 Example of the pre-constructed participant identity database

[0031] In the above Table 1, the identity number is the unique identity identifier assigned by the system, which is used to index database records; the name represents the real name of the attendee; the position corresponds to the name of the position to which the attendee belongs and is used for subsequent permission association; the feature vector encoding is the standard face feature vector extracted from the captured image during the registration phase, and the format is a 128-dimensional or 256-dimensional real number vector; the feature encryption value is the ciphertext obtained by encrypting the feature vector encoding F to ensure data security and is used for privacy protection during storage and transmission; the registration time represents the timestamp when the attendee's identity information is completed and written into the database.

[0032] S4 specifically includes: S41: Receive the identity label data or temporary identity label data output by S3, and retrieve the corresponding identity attribute information according to the identity number in the identity label data or the temporary number in the temporary identity label data; S42: Based on the preset permission rule table in the conference management system, use the position information in the identity attribute information retrieved in S41 as an index to automatically determine the permission level of the attendee; S43: Determine whether the current attendee identity data comes from temporary identity label data. If the attendee data comes from temporary identity label data, add an abnormal status flag for the attendee; if it comes from the registered identity label data, set the abnormal status flag to the default normal status; S44: Merge the name and position information obtained in S41, the permission level obtained in S42, and the abnormal status flag determined in S43 to form a structured data record and generate composite tag information for subsequent display; the above steps ensure the accuracy of the content and the clarity of the structure of the composite tag information of the attendee generated by the conference management system through a clear identity information retrieval, automatic mapping of permission rules, and abnormal status judgment process, can identify unregistered attendees in real time, and enhance the security and reliability of the identity management at the conference site.

[0033] Example of Preset Permission Rules in Table 2

[0034] In the above Table 2, the position category represents the position information marked by the attendee during registration and is used as the basis for permission grading judgment; the permission level is the permission control code defined inside the conference system, and the levels are L5 to L1 from high to low; the permission description is a detailed description of the operation capabilities of each level of permission and is used for the permission execution logic mapping of the conference management system.

[0035] When the participant position retrieved in S41 is the department head, the system will search for a matching item in the permission rule table and assign the participant the L3 permission level, with the permission description being the right to speak normally and participate in the agenda; if it is a temporary numbered identity tag, it will default to the category of temporary visitor / unregistered person, be assigned the L1 permission level, and an exception status flag will be set in the subsequent steps.

[0036] S5 specifically includes: S51: Receive the composite tag information generated in S4, perform homomorphic encryption processing using a preset public key, encrypt each data field containing name, position, permission level, and exception status flag item by item, and generate corresponding ciphertexts. S52: After concatenating the ciphertexts of each field generated in S51 in a predetermined order, add a predefined data packet header and a checksum tail marker, and encapsulate them into a complete encrypted tag data packet that conforms to the parsing protocol of the conference management terminal. S53: Send the encrypted tag data packet encapsulated in S52 to the conference management terminal through a preset independent communication channel using a secure communication protocol. This communication channel is physically isolated from the video stream communication channel and is only used to transmit encrypted identity tag information data. S54: After receiving the encrypted tag data packet, the conference management terminal checks the header and tail of the data packet according to the agreed parsing protocol. After passing the check, it enters the subsequent decryption step; the above steps significantly improve the security and privacy protection level during the data transmission process by performing homomorphic encryption on the composite tag information and transmitting it independently and securely, preventing information leakage and tampering, and ensuring the security and integrity of the identity data in the conference scenario.

[0037] S6 specifically includes: S61: After the conference management terminal receives the encrypted tag data packet transmitted in S5, perform homomorphic decryption on the ciphertexts of each field in the data packet item by item using the pre-stored private key, and decrypt to obtain the plaintext data of the composite tag information. S62: Automatically associate the decrypted composite tag information with the corresponding participant image position coordinates in the real-time video stream of the participant according to the participant identity number or temporary number. After performing image coordinate matching, display the name, position, permission level, and exception status flag near the corresponding participant position. S63: Real-time monitor the position information of the participant. If it is detected that the position coordinates of the participant have changed compared to the previous moment's position, trigger a position movement event, recalculate the display coordinates of the tag information in the video frame, and immediately update the overlay position. S64: Real-time monitor the status information of the privilege levels of participants in the conference management system. If the system detects a change in the privilege level of a participant, update the data content of the composite marker information according to the latest privilege level, and synchronously refresh the overlay marker displayed on the video screen; S65: Continuously execute S63 and S64 to ensure that the marker information of the participants displayed in the real-time video stream is always consistent with the actual positions and real-time privilege statuses of the participants. Through the real-time decryption of the encrypted marker data, precise video position matching, and dynamic monitoring of position movement and privilege changes, the above steps achieve the real-time, accurate, and dynamic update of the participant identity marker information, improve the real-time performance and accuracy of the participation management process, avoid information delay or misaligned display, and ensure efficient and orderly on-site management.

[0038] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of the present invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0039] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An automatic marking method for intelligent conference participants based on image recognition, characterized in that, It includes the following steps: S1: Through multi-angle image acquisition devices deployed in the meeting scenario, the facial images of the participants are obtained in real time, and the facial images are subjected to light compensation, pose correction, and occlusion area segmentation processing to generate structured participant image data; S2: Extract the facial features of the participants from the structured participant image data. For the scenarios of wearing masks and glasses occlusion, a local feature enhancement algorithm is used to extract the geometric distribution features of the eye area and the texture features of the nose bridge area to generate an anti-occlusion feature vector; S3: Match the anti-occlusion feature vector with a pre-constructed participant identity database, which includes the encrypted biometric codes, position information, and permission labels of the registered participants; If the match is successful, the corresponding identity label data is output; If the match fails, an exception mark for unregistered personnel is triggered, and temporary identity label data including a temporary number and the admission time is generated; S4: According to the identity label data or the temporary identity label data, associate the preset permission rules in the conference management system to generate composite label information including name, position, permission level, and exception status identifier; S5: Perform homomorphic encryption processing on the composite label information to generate an encrypted label data packet; S6: In the interactive interface of the conference management terminal, decrypt the encrypted label data packet and superimpose it on the real-time video stream of the participants, and trigger the dynamic refresh and repositioning of the label information based on the events of the participant's position movement or permission change.

2. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, characterized in that The specific content of S1 includes: S11: Through the camera groups preset at the four corners and the central position in the meeting scenario, synchronously trigger shooting to cover the full-angle facial area of the participants, and generate an original multi-view image set; S12: Perform light intensity analysis on the original multi-view image set. For the overexposed or underexposed areas in a single-frame image, a multi-frame fusion compensation algorithm is used to extract the optimal light pixels of each area to generate a light-balanced image; S13: Based on the light-balanced image, detect the facial key point coordinates, and determine the head deflection angle by calculating the spatial vector angle between the tip of the nose point and the points of the two earlobes; when the deflection angle exceeds the preset threshold, use affine transformation to correct the facial image to the front view angle to generate a pose-normalized image; S14: Perform edge detection on the pose-normalized image, identify the contour of the occlusion area of the mask, glasses or hand, and generate a repaired and filled image of the occlusion area according to the facial skin texture features that are not occluded; S15: Bind the repaired and filled image with the associated acquisition timestamp, camera position code, light compensation parameter, and pose correction parameter, and encapsulate it as structured participant image data containing multi-dimensional metadata.

3. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, wherein The specific content of S2 includes: S21: Receive the structured participant image data generated in S1, locate the facial area in the image, and use a preset facial key point detection model to mark the key points such as the corners of the eyes, the ends of the eyebrows, the center of the eyebrows, the tip of the nose, and the root of the nose, and delimit the boundary range of the eye area and the nose bridge area; S22: In the delimited eye area, calculate the geometric relationships between the key points, including the Euclidean distance and the relative angle between the end of the eyebrow and the center of the eyebrow, and between the corners of the eyes, to construct the local geometric distribution features of the face; S23: Within the nasal bridge region, use the local texture gradient analysis method to analyze the gray-scale change distribution in this region, extract the gray-scale gradient information characterizing the skin texture difference, and extract multi-scale local texture features in the way of sliding a fixed window. S24: Stitch and fuse the geometric distribution features and texture features obtained in S22 and S23 to generate a one-dimensional anti-occlusion feature vector.

4. The automatic labeling method for intelligent conference participants based on image recognition according to claim 3, wherein The specific content of S21 includes: S211: Based on the structured participant image data, obtain the image coordinates of the eye corner points, the end points of the eyebrows, the center point of the eyebrows, the tip of the nose point, and the root of the nose point through a pre-trained face key point detection model. S212: According to the detected key point coordinates, define the bounding box of the periorbital region. Taking the outer eye corner points of the left and right eyes as the center, the width of the box body is taken as 1.5 times the Euclidean distance between the two outer eye corner points, the height is taken as 0.8 times the vertical distance from the center point of the eyebrows to the tip of the nose point, and the center line of the box body coincides with the line connecting the two outer eye corners. S213: Define the bounding box of the nasal bridge region. Taking the midpoint between the root of the nose point and the tip of the nose point as the center of the box body, the height of the box body is set as 1.2 times the Euclidean distance between the root of the nose point and the tip of the nose point, the width of the box body is 0.6 times the horizontal distance between the two inner eye corner points, and the vertical center line of the box body coincides with the line connecting the root of the nose point and the tip of the nose point.

5. The automatic marking method for intelligent conference participants based on image recognition according to claim 3, characterized in that, The specific content of S22 includes: S221: Using the facial image coordinates obtained in S21, taking the center point of the eyebrows as the reference, calculate the Euclidean distances between the center point of the eyebrows and the end point of the left eyebrow, the center point of the eyebrows and the end point of the right eyebrow, and the left outer eye corner point and the right outer eye corner point respectively to obtain the characteristic lengths of the periorbital region. S222: Calculate the angles formed by connecting the center point of the eyebrows with the end points of the left and right eyebrows respectively, and the angles formed by connecting the center point of the eyebrows with the left and right outer eye corner points respectively to obtain the characteristic angles of the periorbital region. S223: Normalize the Euclidean distances obtained in S221 and the angles obtained in S222. S224: Sequentially stitch the normalized characteristic lengths and characteristic angles to form the facial local geometric distribution features.

6. The automatic marking method for intelligent conference participants based on image recognition according to claim 3, characterized in that, The specific content of S23 includes: S231: Extract the gray-scale image sub-region from the bounding box of the nasal bridge region delimited in S21, denoted as the region to be analyzed. S232: Taking each pixel point in the region to be analyzed as the center, calculate the gradient amplitude of the gray-scale value along the horizontal and vertical directions of the image respectively to obtain the local gradient intensity information. S233: Based on the gradient amplitude, use a square window of a preset size to slide and traverse the region to be analyzed at a fixed step length. The histogram distribution of the gradient amplitude is statistically calculated within each window to form the local texture feature histogram. S234: Change the window size, repeat S233, extract the gradient histograms with different window sizes respectively, and sequentially stitch the histogram features under different window sizes to form the local texture features with multi-scale information.

7. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, characterized in that, The specific content of S3 includes: S31: Extract the encrypted biometric feature codes of the registered participants from the pre-constructed participant identity database, and decrypt the encrypted biometric feature codes and convert them into identity feature vectors in the standard format. S32: Calculate the similarity between the anti-occlusion feature vector obtained in S2 and each identity feature vector one by one, and obtain the matching similarity values between each identity feature vector and the anti-occlusion feature vector through the cosine similarity method; S33: Sort the matching similarity values obtained in S32, select the identity feature vector with the highest matching similarity, and determine whether the similarity value is higher than the preset identity matching threshold; S34: If the highest matching similarity value in S33 exceeds the identity matching threshold, it is determined that the match is successful, and the identity label data associated with the corresponding identity feature vector is output; if the highest matching similarity value does not exceed the identity matching threshold, it is determined that the match fails, and the unregistered person exception marking program is triggered to output the temporary identity label data including the temporary number and the entry time.

8. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, characterized in that, The specific steps of S4 are as follows: S41: Receive the identity label data or the temporary identity label data output by S3, and retrieve the corresponding identity attribute information according to the identity number in the identity label data or the temporary number in the temporary identity label data; S42: Based on the preset permission rule table in the conference management system, use the position information in the identity attribute information retrieved in S41 as an index to automatically determine the permission level of the attendee; S43: Determine whether the current attendee identity data comes from the temporary identity label data. If the attendee data comes from the temporary identity label data, add an exception status flag for the attendee; if it comes from the registered identity label data, set the exception status flag to the default normal status; S44: Merge the name and position information obtained in S41, the permission level obtained in S42, and the exception status flag determined in S43 to form a structured data record and generate composite marking information.

9. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, characterized in that The specific steps of S5 are as follows: S51: Receive the composite marking information generated by S4, perform homomorphic encryption processing using the preset public key, and encrypt each data field including the name, position, permission level, and exception status flag item by item to generate the corresponding ciphertext; S52: After splicing the ciphertexts of each field generated in S51 in a predetermined order, add a predefined data packet header and a checksum tail mark, and encapsulate them into a complete encrypted marking data packet that conforms to the parsing protocol of the conference management terminal; S53: Through a preset independent communication channel, send the encrypted marking data packet encapsulated in S52 to the conference management terminal using a secure communication protocol. This communication channel is physically isolated from the video stream communication channel and is only used to transmit encrypted identity marking information data; S54: After receiving the encrypted marking data packet, the conference management terminal checks the header and tail of the data packet according to the agreed parsing protocol.

10. The automatic marking method for intelligent conference participants based on image recognition according to claim 1, wherein, The specific steps of S6 are as follows: S61: After receiving the encrypted marking data packet transmitted by S5, the conference management terminal performs homomorphic decryption on the ciphertexts of each field in the data packet item by item using the pre-stored private key to decrypt and obtain the plaintext data of the composite marking information; S62: Automatically associate the decrypted composite tag information with the corresponding participant image position coordinates in the real-time video stream of the participant according to the participant identity number or the temporary number. After performing image coordinate matching, display the name, position, permission level, and anomaly status identifier near the corresponding participant position; S63: Continuously monitor the position information of the participant. If it is detected that the position coordinates of the participant have changed compared to the previous moment, trigger a position movement event, recalculate the display coordinates of the tag information in the video frame, and immediately update the overlay position; S64: Continuously monitor the status information of the participant permission level in the conference management system. If the system detects that the permission level of the participant has changed, update the data content of the composite tag information according to the latest permission level, and synchronously refresh the overlay tag displayed on the video frame; S65: Continuously execute S63 and S64 to ensure that the participant tag information displayed in the real-time video stream is always consistent with the actual position and real-time permission status of the participant.

Citation Information

Patent Citations

  • Conference management method and system based on face recognition and readable storage medium

    CN110072075A

  • Encryption method and system of LED conference display screen

    CN119167391A

  • Indoor safety monitoring system and method based on image recognition

    CN119854458A

  • Human eye model training method, human eye recognition method, apparatus, device and medium

    WO2019232866A1