Conference participant information management method and system
By collecting face images of multiple preset expressions of participants and audio information of preset statements, using the long and short-term memory network to extract audio features, calculate voiceprint features, and updating face features through a combination scheme, the problem of poor information management of participants in the existing technology is solved, and efficient identity authentication and information management is achieved.
Patent Information
- Application Number
- CN202510474214.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When the prior art realizes identity authentication for participants, it is impossible to effectively manage participants’ information, resulting in low accuracy of identity authentication and high storage pressure.
By collecting face images of multiple preset expressions of participants and audio information of preset statements, the audio features are extracted using the long and short-term memory network, the voiceprint features are calculated, and the face features are updated through the combination scheme to achieve identity authentication.
While ensuring the accuracy of identity authentication, it reduces the pressure of information storage and realizes effective management of information of participants.
Smart Images

Figure CN119989322A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of information management, and in particular to a method and system for managing information of meeting participants. Background Art
[0002] With the rapid development of Internet technology, online meetings are gradually replacing offline meetings with their convenience and speed. During online meetings, participants need to click on the meeting link on the client and authenticate their identities. Only when the authentication is successful can they enter the online meeting to ensure the security of the online meeting. In order to realize the identity authentication of the participants, the participant information of the parameter participants needs to be entered in advance. How to manage the participant information to ensure the accuracy of the identity authentication when entering the online meeting is an urgent problem to be solved.
[0003] At present, the patent application document with application publication number CN113420274A discloses a participant access management system and method based on trusted identity authentication, wherein the method includes: identifying certificate information and matching and verifying it with the certificate registration information in the trusted identity authentication platform; prompting the participants that they are about to perform electronic signatures and face verification, and collecting the posture and voice information of the participants for comprehensive judgment; reading third-party voice and retinal biometric verification information based on the certificate information and storing it in a secure cache memory, while reading the voice and retinal data of the participants for information matching and verification, and realizing trusted identity authentication by combining partial face matching and full face matching.
[0004] The above method integrates multiple information such as the posture, voice information, retinal data and facial information of the participants to achieve trusted identity authentication. However, not all data is valid for identity authentication. Integrating multiple information will not only not increase the accuracy of identity authentication, but will increase storage pressure and fail to achieve effective management of participant information. Summary of the invention
[0005] In order to solve the technical problem of being unable to effectively manage attendee information, the present application provides a method and system for managing attendee information, which can screen attendee information, reduce information storage pressure while ensuring the accuracy of identity authentication, and achieve effective management of attendee information.
[0006] In a first aspect, the present application provides a method for managing information of participants, the method comprising: collecting facial images of multiple preset expressions and audio information of preset sentences of any participant; dividing the audio information into audio segments of each character in the preset sentence, obtaining the audio features of each character using a long short-term memory network, weightedly summing the audio features of each character according to the variance of the audio features of each participant, and obtaining the voiceprint features of the participant; performing feature extraction on the facial images to obtain the facial features of each preset expression, initializing a combination scheme of the preset expressions of each participant, and calculating the average facial features of the selected preset expressions according to the average facial features of the selected preset expressions in the combination scheme. Calculate the global significance of the combination scheme and the individual significance of each participant, the individual significance is the minimum difference between the average facial features of the participant and other participants, and the variance of the global significance is negatively correlated with the individual significance; take the sum of the minimum values of the global significance and the individual significance as the objective function, update the combination scheme of each participant, and take the combination scheme corresponding to the maximum value of the objective function as the target scheme; store the voiceprint features of each participant, the target expression in the target scheme and the average facial features of each target expression to realize the information management of the participants, and the average facial features and voiceprint features of each target expression are used for identity authentication.
[0007] Collect facial images of multiple preset expressions and audio information of preset sentences of any participant, the preset sentence includes multiple characters, and use the long short-term memory network to extract the audio features of each character in the audio information. Because different participants will have different pronunciations when reading the same character, when all participants are reading a character, the greater the difference between the participants, it means that the audio feature corresponding to the character can accurately reflect the voiceprint features of each participant. Therefore, the audio features of each character are weighted and summed according to the variance of the audio features of each participant to obtain the voiceprint features of each participant; further, obtain the facial features of each preset expression of a participant, initialize the combination scheme of the preset expressions of each participant, each participant corresponds to a combination scheme, and calculate the global significance and The individual significance of each participant can measure the difficulty of face recognition of the corresponding participant, and the global significance can measure the degree of difference between the individual significance of all participants. The sum of the global significance and the minimum individual significance is used as the objective function, and the combination plan of each participant is continuously updated until the target plan of each participant is obtained when the objective function reaches the maximum value. The identity authentication of the corresponding participant can be accurately realized according to the facial features corresponding to the target expression in the target plan; the voiceprint features of each participant, the target expression in the target plan and the average facial features of each target expression are stored, and identity authentication is performed based on the voiceprint features and the average facial features of each target expression. This can reduce the storage pressure while ensuring the accuracy of identity authentication, and realize the effective management of participant information.
[0008] Preferably, dividing the audio information into audio segments of each character in the preset sentence includes: performing frame processing on the audio information and calculating the short-time energy of any timestamp to obtain a short-time energy sequence, dividing the short-time energy sequence according to extreme points to obtain multiple time intervals; in response to the number of time intervals being equal to the number of characters in the preset sentence, the audio information in each time interval is used as the audio segment of the corresponding character, otherwise, the audio information of the participants is re-collected.
[0009] When the number of time intervals is equal to the number of characters in the preset sentence, a one-to-one correspondence is achieved between the time intervals and the characters; when the number of time intervals is not equal to the number of characters in the preset sentence, it indicates that missing words or repeated readings have occurred in the audio information. At this time, it is necessary to re-collect the audio information of the participant to ensure that an audio clip of each character of the participant can be obtained.
[0010] Preferably, obtaining the audio features of each character using the long short-term memory network includes: inputting the audio information into the long short-term memory network to obtain the short-term vector of each timestamp; and taking the average value of all short-term vectors of timestamps in the audio segment of any character as the audio feature of the character.
[0011] Preferably, participants Voiceprint features for: , For participants In character The audio characteristics of Characters for all participants The audio feature variance is is the sum of the audio feature variances of all characters of all participants, The number of characters in the preset sentence.
[0012] Different participants will have different pronunciations when reading the same character. When all participants are reading one character, the greater the difference between the participants, the more likely it is that the audio features corresponding to the character can accurately reflect the voiceprint features of each participant. Therefore, the audio features of each character are weighted and summed according to the variance of the audio features of each participant to accurately obtain the voiceprint features of each participant.
[0013] Preferably, the combination scheme is a multi-dimensional 01 vector, and when the value of any dimension is 1, it indicates that the preset expression corresponding to the dimension is selected, otherwise, the preset expression corresponding to the dimension is not selected.
[0014] Preferably, extracting features from the face image includes: extracting features from the face image based on an autoencoder network or a principal component analysis algorithm.
[0015] Preferably, participants The calculation process of the individual significance of the participants is as follows: The Euclidean distance between the average facial features of other participants, and the minimum Euclidean distance is taken as the participant individual significance.
[0016] The individual significance minimum value can reflect the highest difficulty of identity authentication for all participants. The larger the individual significance minimum value, the smaller the maximum difficulty, which means that the difficulty for all participants to complete identity authentication is smaller, thus achieving accurate quantification of the difficulty of identity authentication.
[0017] Preferably, global saliency Satisfies the relationship: , is the variance of the individual significance.
[0018] Preferably, the identity authentication process includes: in response to any participant receiving an invitation to join the meeting, issuing an action instruction to the participant, the action instruction including the target expression of the participant; collecting real-time voiceprint features and real-time facial features of each target expression, and calculating the real-time average facial features; calculating the facial similarity between the real-time average facial features and the average facial features, and the voiceprint similarity between the real-time voiceprint features and the voiceprint features; in response to the facial similarity and the voiceprint similarity being greater than the preset similarity, the identity authentication is successful; otherwise, the identity authentication fails.
[0019] Since the target plans of each participant are different, the action instructions of each participant are also different. During the identity authentication process, it is only necessary to collect face images under the target expression, avoiding the collection of face images under all preset expressions, improving authentication efficiency, and reducing the amount of calculation during identity authentication.
[0020] In a second aspect of the present application, there is also provided a participant information management system, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a participant information management method according to the first aspect of the present application is implemented.
[0021] The technical solution of this application has the following beneficial technical effects: First, facial images of multiple preset expressions and audio information of preset sentences of any participant are collected. The preset sentence includes multiple characters, and the audio features of each character in the audio information are extracted using the long short-term memory network. Since different participants have different pronunciations when reading the same character, when all participants are reading a character, the greater the difference between the participants, the more accurate the audio feature corresponding to the character can be. The audio feature of each character is weighted and summed according to the variance of the audio features of each participant to obtain the voiceprint features of each participant; further, the facial features of each preset expression of a participant are obtained, and the combination scheme of the preset expressions of each participant is initialized. Each participant corresponds to a combination scheme, and the global significance of the combination scheme is calculated. and the individual significance of each participant. The individual significance can measure the difficulty of face recognition of the corresponding participant. The global significance can measure the degree of difference between the individual significances of all participants. The sum of the global significance and the minimum individual significance is used as the objective function. The combination scheme of each participant is continuously updated until the target scheme of each participant is obtained when the objective function reaches the maximum value. The identity authentication of the corresponding participant can be accurately realized according to the facial features corresponding to the target expression in the target scheme; the voiceprint features of each participant, the target expression in the target scheme and the average facial features of each target expression are stored. Identity authentication is performed based on the voiceprint features and the average facial features of each target expression. This can reduce the storage pressure while ensuring the accuracy of identity authentication, and realize the effective management of participant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of a method for managing participant information according to an embodiment of the present application.
[0023] Figure 2 It is a structural block diagram of a participant information management system according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0025] According to the first aspect of the present application, the present application provides a method for managing participant information, which is used to screen and store the participant information of all participants so that participants can quickly and accurately complete identity authentication when joining an online meeting.
[0026] Figure 1FIG. 1 is a flow chart of a method for managing information of meeting participants according to an embodiment of the present application. Figure 1 As shown, the participant information management method includes steps S101 to S105, which are described in detail below.
[0027] S101, collecting facial images of multiple preset expressions and audio information of preset sentences of any participant.
[0028] In one embodiment, the participant is any participant of the online conference system. In an enterprise, the participant includes all employees in the enterprise; the preset expressions include multiple expressions such as opening the mouth, blinking, smiling or raising eyebrows, and the facial image of any participant under each preset expression is collected, and the audio information of the participant when reading the preset sentence is collected.
[0029] Among them, the preset sentence is set in advance to reflect the voiceprint characteristics of the participants and is not limited here.
[0030] In this way, the number of preset expressions is recorded as N, and a participant can collect N facial images and a string of audio information. The N facial images and a string of audio information are used to realize identity authentication when joining an online meeting. However, directly storing the N facial images and audio information of all participants will take up a lot of side storage resources. Therefore, it is necessary to screen the information of these participants.
[0031] S102, dividing the audio information into audio segments of each character in a preset sentence, using a long short-term memory network to obtain the audio features of each character, and performing weighted summation of the audio features of each character according to the variance of the audio features of each participant to obtain the voiceprint features of the participant.
[0032] In one embodiment, the preset sentence includes multiple characters, and the audio information is divided to obtain an audio segment corresponding to each character. Specifically, dividing the audio information into audio segments of each character in the preset sentence includes: performing frame processing on the audio information and calculating the short-time energy of any timestamp to obtain a short-time energy sequence, dividing the short-time energy sequence according to extreme points to obtain multiple time intervals; in response to the number of time intervals being equal to the number of characters in the preset sentence, the audio information in each time interval is used as the audio segment of the corresponding character, otherwise, the audio information of the participants is re-collected.
[0033] The framing window of any timestamp is a time period centered on the timestamp and including multiple timestamps on both sides. Framing the audio information is the process of obtaining each timestamp framing window.
[0034] Short-time energy is used to measure the energy intensity of an audio signal within a time window. It is a common technology in the field of speech signal processing and will not be described in detail here. When the number of time intervals is equal to the number of characters in the preset sentence, a one-to-one correspondence between time intervals and characters is achieved; when the number of time intervals is not equal to the number of characters in the preset sentence, it indicates that there are missing words or repeated readings in the audio information. At this time, the audio information of the participant needs to be re-collected to ensure that the audio clip of each character of the participant can be obtained.
[0035] In one embodiment, obtaining the audio features of each character using a long short-term memory network includes: inputting audio information into the long short-term memory network to obtain a short-term vector of each timestamp; and taking the average value of all short-term vectors of timestamps in an audio segment of any character as the audio feature of the character.
[0036] Among them, the audio information is a time series data, and the long short-term memory network can extract the long-time vector and short-time vector of each time stamp in the time series data. The long-time vector is used to characterize the time series characteristics of all time series data from the beginning of the time series data to the corresponding time stamp, and the short-time vector is used to characterize the time series characteristics of the corresponding time stamp in the time series data. It should be noted that the training process of the long short-term memory network is a well-known technology for those skilled in the art and will not be repeated here.
[0037] Since different participants may pronounce the same character differently, when all participants read a character, the greater the difference between the participants, the more likely it is that the audio features corresponding to the character can accurately reflect the voiceprint features of each participant. Therefore, the voiceprint features of each participant are calculated based on the variance of the audio features of each participant.
[0038] Specifically, the participants Voiceprint features for: , For participants In character The audio characteristics of Characters for all participants The audio feature variance is is the sum of the audio feature variances of all characters of all participants, The number of characters in the preset sentence.
[0039] In this way, the voiceprint characteristics of each participant are obtained.
[0040] S103, extracting features from the facial image to obtain facial features of each preset expression, initializing a combination scheme of preset expressions of each participant, and calculating the global significance of the combination scheme and the individual significance of each participant based on the average facial features of the selected preset expressions in the combination scheme, wherein the individual significance is the minimum difference between the average facial features of the participant and other participants, and the variance of the global significance is negatively correlated with the individual significance.
[0041] In one embodiment, feature extraction of facial images can be performed based on an autoencoder network or a principal component analysis algorithm to obtain facial features corresponding to each facial image. Because a participant can collect facial images with multiple preset expressions, the facial features of the participant under each preset expression can be obtained.
[0042] The combination scheme is a multi-dimensional 01 vector. When the value of any dimension is 1, it indicates that the preset expression corresponding to the dimension is selected. Otherwise, the preset expression corresponding to the dimension is not selected. If the number of preset expressions is N, the dimension of the combination scheme is also N. One dimension in the combination scheme corresponds to one preset expression. For example, if the number of preset expressions is 5, the combination scheme is , it means selecting the facial features of the first preset expression and the fourth preset expression.
[0043] In one embodiment, a combination scheme of each participant is initialized, and each combination scheme corresponds to a selection of a preset expression. Under each combination scheme of each participant, the global saliency and the individual saliency of each participant are calculated. The individual saliency is the minimum difference in facial features between the participant and other participants. For example, participants The greater the minimum difference in facial features between the participant and other participants, the The more prominent the facial features are, the more beneficial it is for the participants. Identification of participants The greater the individual significance; the global significance is used to evaluate the degree of difference between the individual significances of all participants. When the individual significances of all participants are basically consistent, all participants can obtain accurate identity recognition results. At this time, the global significance is greater.
[0044] Specifically, the participants Average facial features for: , For participants In the combination of The facial features of the selected preset expressions, For participants The number of preset expressions selected in the combination scheme.
[0045] After obtaining the average facial features of all participants, the individual significance of each participant can be calculated. The calculation process of the individual significance of the participants is as follows: The Euclidean distance between the average facial features of other participants, and the minimum Euclidean distance is taken as the participant individual significance.
[0046] The global significance is negatively correlated with the variance of individual significance. Satisfies the relationship: , is the variance of the individual significance.
[0047] In this way, the combination scheme of all participants is initialized, which includes the selection of preset expressions by each participant. The combination scheme of each participant is different, and the individual saliency of each participant is calculated under the combination scheme of all participants. The individual saliency is used to evaluate the difficulty of identity authentication of each participant; at the same time, the global saliency of all participants is calculated to evaluate the degree of difference between the individual saliencies of all participants.
[0048] S104, taking the sum of the global significance and the minimum individual significance as the objective function, updating the combination plan of each participant, and taking the combination plan corresponding to the maximum value of the objective function as the target plan.
[0049] In one embodiment, the individual significance minimum value can reflect the highest difficulty of identity authentication for all participants. The larger the individual significance minimum value is, the smaller the highest difficulty is, which means that the difficulty for all participants to complete identity authentication is smaller. Therefore, the individual significance minimum value is used as part of the objective function. Furthermore, the larger the global significance is, the more difficult it is for all participants to complete identity authentication. The sum of the individual significance minimum value and the global significance is used as the objective function. When the value of the objective function is larger, it means that the difficulty for all participants to complete identity authentication is at an approximately smaller value, and all participants can obtain accurate identity authentication results.
[0050] After defining the objective function, the maximum value of the objective function is taken as the optimization goal, and the combination plan of each participant is continuously updated until the maximum value of the objective function is reached. The combination plan corresponding to each participant at this time is used as the target plan of the corresponding participant. Specifically, the particle swarm optimization algorithm, simulated annealing algorithm or genetic algorithm can be used to update the combination plan of each participant until the target plan of each participant is obtained.
[0051] In this way, a target scheme for each participant is determined. A target scheme for a participant includes at least one target expression. The identity authentication of the corresponding participant can be accurately achieved according to the facial features corresponding to the target expression.
[0052] S105, storing the voiceprint features of each participant, the target expression in the target solution and the average facial features of each target expression, to achieve information management of the participants, and the average facial features and voiceprint features of each target expression are used for identity authentication.
[0053] In one embodiment, one is used to correspond to a target scheme, and a target scheme includes at least one target expression. In order to ensure the accuracy of identity authentication while reducing the pressure of information storage, there is no need to store facial images of all preset expressions and audio information of preset sentences. It is only necessary to store the voiceprint features, target expressions and the average facial features of each target expression of the participants to realize the information management of the participants. At the same time, the voiceprint features and the average facial features of each target expression are used to integrate the two dimensions of image features and audio features to realize identity authentication when joining an online meeting.
[0054] Specifically, the identity authentication process includes: in response to any participant receiving an invitation to join the meeting, issuing an action instruction to the participant, the action instruction including the target expression of the participant; collecting real-time voiceprint features and real-time facial features of each target expression, and calculating the real-time average facial features; calculating the facial similarity between the real-time average facial features and the average facial features, and the voiceprint similarity between the real-time voiceprint features and the voiceprint features. In response to the facial similarity and the voiceprint similarity being greater than the preset similarity, the identity authentication is successful; otherwise, the identity authentication fails.
[0055] Among them, the preset similarity is 0.8.
[0056] It is understandable that since the target plans of each participant are different, the action instructions of each participant are also different. During the identity authentication process, it is only necessary to collect facial images under the target expression, avoiding the collection of facial images under all preset expressions, thereby improving authentication efficiency and reducing the amount of calculation during identity authentication.
[0057] According to the second aspect of the present application, the present application also provides a participant information management system. Figure 2 is a structural block diagram of a participant information management system according to an embodiment of the present application. Figure 2As shown, the system 50 includes a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a method for managing participant information according to the first aspect of the present application is implemented. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, whose settings and functions are known in the art, and therefore will not be described in detail here.
[0058] It should be pointed out that, for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present application, which all fall within the scope of protection of the present application.
Claims
1. A method for managing information of meeting participants, characterized in that: The management method comprises: Collect facial images of multiple preset expressions and audio information of preset sentences of any participant; Divide the audio information into audio segments of each character in the preset sentence, use the long short-term memory network to obtain the audio features of each character, and perform weighted summation of the audio features of each character according to the variance of the audio features of each participant to obtain the voiceprint features of the participant; Perform feature extraction on the face image to obtain the face features of each preset expression, initialize the combination scheme of the preset expressions of each participant, and calculate the global significance of the combination scheme and the individual significance of each participant according to the average face features of the selected preset expressions in the combination scheme, wherein the individual significance is the minimum difference between the average face features of the participant and other participants, and the variance of the global significance is negatively correlated with the individual significance; Taking the sum of the global significance and the minimum individual significance as the objective function, update the combination plan of each participant, and take the combination plan corresponding to the maximum value of the objective function as the target plan; The voiceprint features of each participant, the target expression in the target plan and the average facial features of each target expression are stored to achieve information management of the participants. The average facial features and voiceprint features of each target expression are used for identity authentication.
2. A method for managing information of participants according to claim 1, characterized in that: The audio segments that divide the audio information into the characters in the preset sentence include: The audio information is framed and the short-time energy of any time stamp is calculated to obtain a short-time energy sequence, and the short-time energy sequence is divided according to the extreme value points to obtain multiple time intervals; In response to the number of the time intervals being equal to the number of characters in the preset sentence, the audio information in each time interval is used as the audio segment of the corresponding character; otherwise, the audio information of the conference participants is recollected.
3. A method for managing information of participants according to claim 1, characterized in that: The audio features of each character obtained using the long short-term memory network include: Input the audio information into the long short-term memory network to obtain the short-term vector of each time stamp; The average value of all time stamp short-time vectors in the audio segment of any character is taken as the audio feature of the character.
4. A method for managing information of participants according to claim 1, characterized in that: Participants Voiceprint features for: , For participants In character The audio characteristics of Characters for all participants The audio feature variance is is the sum of the audio feature variances of all characters of all participants, The number of characters in the preset sentence.
5. A method for managing information of participants according to claim 1, characterized in that: The combination scheme is a multi-dimensional 01 vector. When the value of any dimension is 1, it indicates that the preset expression corresponding to the dimension is selected. Otherwise, the preset expression corresponding to the dimension is not selected.
6. A method for managing information of participants according to claim 1, characterized in that: Extracting features from facial images includes: extracting features from facial images based on an autoencoder network or a principal component analysis algorithm.
7. A method for managing information of participants according to claim 1, characterized in that: Participants The calculation process of the individual significance of the participants is as follows: The Euclidean distance between the average facial features of other participants, and the minimum Euclidean distance is taken as the participant individual significance.
8. A method for managing information of participants according to claim 1, characterized in that: Global saliency Satisfies the relationship: , is the variance of the individual significance.
9. A method for managing participant information according to any one of claims 1 to 8, characterized in that: The identity authentication process includes: In response to any participant receiving a meeting invitation, issuing an action instruction to the participant, wherein the action instruction includes a target expression of the participant; Collect real-time voiceprint features and real-time facial features of each target expression, and calculate real-time average facial features; The face similarity between the real-time average face feature and the average face feature, and the voiceprint similarity between the real-time voiceprint feature and the voiceprint feature are calculated. If both the face similarity and the voiceprint similarity are greater than the preset similarity, the identity authentication is successful. Otherwise, the identity authentication fails.
10. A participant information management system, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a method for managing participant information according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
User access management system and method based on trusted identity authentication
CN113420274A
Identity authentication method of network safety access of power system
CN105022994A
Audio-visual conversion control method and system based on deep learning, and storage medium
CN116781856A
Conference information recording method and apparatus, computer device, and storage medium
WO2019227579A1