Conference sign-in method and system based on voiceprint recognition

By dynamically adjusting the arrangement method and scale of the microphone array in the voiceprint recognition conference check-in system, the problem of degradation in acquisition quality caused by fixed configuration is solved, and the system's identification accuracy and reliability in different environments is improved.

CN119724195BActive Publication Date: 2025-08-15GUANGZHOU HUASHAN ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411676389.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-08-15
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

The existing voiceprint recognition conference check-in system cannot adapt to different conference environments due to the fixed microphone configuration, resulting in a decrease in the sound acquisition quality, affecting the recognition accuracy and system reliability.

Method used

By obtaining the first sound source collected by the microphone array in the target area before checking in to the meeting, dynamically adjusting the arrangement method and array size of the microphone array to adapt to the current conference environment, and collecting the second sound source for voiceprint recognition.

Benefits of technology

The collection quality and recognition reliability of the voiceprint recognition conference check-in system in different conference environments has been improved, and the system's ability to adapt to environmental changes has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724195B_ABST
    Figure CN119724195B_ABST
Patent Text Reader

Abstract

A conference check-in method and system based on voiceprint recognition relates to the field of voiceprint recognition. The method comprises: before conference check-in begins, obtaining a first sound source captured by a microphone array in a preset mode within a target area; adjusting the preset mode to a target mode based on the first sound source, wherein the preset mode includes the arrangement and size of the microphone array; and when conference check-in begins, obtaining a second sound source captured by the microphone array in the target mode, obtaining a voiceprint recognition result of the second sound source, and performing conference check-in based on the recognition result. Implementing the technical solution provided in this application can improve the reliability of the voiceprint recognition conference check-in system in different conference environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voiceprint recognition, and specifically to a conference check-in method and system based on voiceprint recognition. Background Art

[0002] As conference management becomes increasingly intelligent, more and more companies and organizations are adopting automated conference attendance systems. In large conferences, with large numbers of attendees and complex, ever-changing environments, quickly and accurately verifying and checking in attendees has become a critical issue. Voiceprint recognition, due to its contactless, efficient, and difficult-to-forge features, is becoming a key technology choice for conference attendance systems.

[0003] Currently, conference check-in systems based on voiceprint recognition typically use fixed microphone arrays to collect voice signals from participants. When participants enter the conference room, these systems collect their voice samples through pre-deployed microphones and match the extracted voiceprint features with samples in a database to achieve identity verification and automatic check-in.

[0004] However, existing voiceprint recognition conference sign-in systems often use fixed microphone configurations that cannot dynamically adjust to the actual meeting environment. When the meeting environment changes, such as the number of participants, the distribution of personnel, or the ambient noise, the fixed microphone configuration may lead to a decline in sound collection quality, affecting the accuracy of voiceprint recognition and thus reducing the reliability of the sign-in system. Summary of the Invention

[0005] The present application provides a conference sign-in method and system based on voiceprint recognition, which can improve the reliability of the voiceprint recognition conference sign-in system in different conference environments.

[0006] In a first aspect of the present application, the present application provides a conference check-in method based on voiceprint recognition, comprising:

[0007] When the meeting check-in has not started, obtaining a first sound source collected by a microphone array in a preset mode in a target area, and adjusting the preset mode to a target mode based on the first sound source, wherein the preset mode includes an arrangement and an array size of the microphone array;

[0008] When the conference sign-in starts, a second sound source collected by the microphone array in the target mode is obtained, a voiceprint recognition result of the second sound source is obtained, and the conference sign-in is performed based on the recognition result.

[0009] In a second aspect of the present application, a conference check-in system based on voiceprint recognition is provided, comprising:

[0010] A microphone array adjustment module is configured to obtain a first sound source collected by a microphone array in a preset mode in a target area when the conference check-in has not yet begun, and adjust the preset mode to a target mode based on the first sound source, wherein the preset mode includes an arrangement and array size of the microphone array;

[0011] The voiceprint recognition sign-in module is used to obtain a second sound source collected by the microphone array in the target mode when the meeting sign-in starts, obtain the voiceprint recognition result of the second sound source, and perform meeting sign-in based on the recognition result.

[0012] In a third aspect of the present application, a computer storage medium is provided, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above method steps.

[0013] In a fourth aspect of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0014] In a fifth aspect of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the method steps described above.

[0015] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0016] By obtaining the first sound source collected by the microphone array in the preset mode in the target area before the start of the meeting check-in, and adjusting the preset mode to the target mode based on the first sound source, the arrangement and array size of the microphone array can adapt to the current meeting environment characteristics, and then at the start of the meeting check-in, the microphone array in the adjusted target mode is used to collect the second sound source and perform voiceprint recognition. Compared with the method of using a fixed microphone configuration scheme in the prior art, this application improves the sound collection quality by dynamically adjusting the microphone array configuration, so that the system can obtain higher-quality voiceprint features, thereby improving the reliability of the voiceprint recognition conference check-in system in different meeting environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of a conference check-in method based on voiceprint recognition provided by an embodiment of the present application;

[0018] Figure 2 This is a structural diagram of a conference check-in system based on voiceprint recognition provided by an embodiment of the present application;

[0019] Figure 3It is a structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0021] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0022] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0023] Please refer to Figure 1 , Figure 1 This is a flow chart of a method for conference attendance based on voiceprint recognition provided by an embodiment of the present application. This method can be implemented by a computer program, a single-chip microcomputer, or run on a conference attendance system based on voiceprint recognition based on a von Neumann architecture. The computer program can be integrated into an application or run as an independent tool application. Specifically, the method may include the following steps:

[0024] Step 101: When the meeting check-in has not started, obtain the first sound source collected by the microphone array in the preset mode in the target area, and adjust the preset mode to the target mode based on the first sound source. The preset mode includes the arrangement and array size of the microphone array.

[0025] The target area refers to the specific space within the conference room used for meeting check-in. Specifically, the target area can be defined as the area surrounded by conference tables, chairs, and podiums, where participants can converse normally and the microphone array can effectively capture sound. For example, in a large conference room, the target area might be the three-meter radius around the conference table; in a multi-purpose conference room, the target area might be the seated audience area; and in a lecture hall, the target area might include the podium and audience seating areas.

[0026] In the embodiment of the present application, the delineation of the target area needs to comprehensively consider the following factors: first, the acoustic propagation characteristics within the area should be relatively stable to avoid severe acoustic attenuation or distortion caused by the spatial structure; second, the area should cover all possible locations of participants who need to sign in; finally, the scope of the area should be within the effective pickup range of the microphone array to ensure the quality of the voiceprint signal collection.

[0027] The target area provides a spatial reference for microphone array deployment and also provides boundary constraints for subsequent area division and sound source acquisition. By properly defining the target area, the system's adaptability to the conference environment can be improved and the accuracy of voiceprint recognition and sign-in can be ensured. In practice, if the target area is too large, the acquisition quality in areas far from the microphone array may be reduced; if the target area is too small, the sign-in audio of some participants may be missed.

[0028] Furthermore, a microphone array refers to an acoustic sensor system consisting of multiple microphone units arranged in a specific geometric pattern within a target area. Specifically, a microphone array can be understood as a combination of at least two microphone units in a predetermined spatial relationship. These microphone units maintain a specific relative position and spacing between them and work together to achieve functions such as sound source localization, sound enhancement, and noise suppression.

[0029] It should be noted that multiple independent microphone units can be set up in the target area, and some of these microphone units can be combined into one or more microphone arrays according to actual needs. For example, a large conference room may have a total of 20 microphone units, of which 8 microphone units can form a linear microphone array for the podium area, and the remaining 12 microphone units can be combined into multiple smaller arrays or remain independent as needed. This flexible configuration allows the system to dynamically adjust the composition of the microphone array based on the specific circumstances of the conference environment.

[0030] In the embodiments of this application, the microphone array is primarily used for collecting and processing high-quality voiceprint information. By working in concert with multiple microphone units, the microphone array enables spatially selective sound collection, effectively suppressing interference from ambient noise and reverberation. Furthermore, by utilizing the time and phase differences between the sound signals collected by each microphone unit in the microphone array, the system can accurately localize and separate sound sources, providing high-quality audio input for subsequent voiceprint recognition.

[0031] Correspondingly, the microphone array in its preset mode captures the sum of the acoustic signals within the target area. The first sound source can be understood as a composite acoustic signal containing multiple acoustic components, including ambient noise, conversations between participants upon entering the room, operating sounds of equipment like air conditioners, and room reverberation. For example, as participants enter the conference room and take their seats or wait in the queue area, the microphone array captures all acoustic signals, including human voices, ambient noise, and the resulting reflections and reverberations within the conference room. These signals are all considered components of the first sound source.

[0032] In the embodiment of the present application, the first sound source is mainly used for the system to pre-evaluate and adapt the conference environment. By analyzing the acoustic characteristics such as the noise level, reverberation characteristics, and sound source distribution in the first sound source, the system can promptly detect whether the microphone array in the preset mode is suitable for the current conference environment. For example, by analyzing the noise components in the first sound source, the signal-to-noise ratio level of the current environment can be evaluated; by analyzing the reverberation characteristics therein, the acoustic characteristics of the conference room can be understood; by analyzing the spatial distribution of the human voice components, the main activity areas of the participants can be determined. This information provides a basis for subsequent adjustments to the arrangement and array size of the microphone array.

[0033] The preset mode, including the arrangement and array size of the microphone array, refers to the initial working configuration parameters of the microphone array before the meeting check-in begins. Among them, the arrangement refers to the geometric distribution of the microphone units in space, and the array size refers to the number of microphone units participating in the array. Specifically, the preset mode can be understood as a set of default parameters pre-set by the system at the factory or by the system administrator. For example, the arrangement can be an evenly spaced linear arrangement with a spacing of 10 cm, and the array size can be a standard configuration consisting of 8 omnidirectional microphone units.

[0034] In specific application scenarios, preset mode arrangements may include, but are not limited to, linear arrangement (microphone units evenly distributed along a straight line), circular arrangement (microphone units arranged in a circle), rectangular arrangement (microphone units arranged in a rectangular array), and random arrangement (irregular arrangement based on the conference room layout). The array size can be preset based on the conference room size and the expected number of attendees, typically with different numbers of microphone units, such as 4, 8, or 16. Preset modes typically select relatively general configuration parameters to accommodate most common meeting scenarios.

[0035] The preset mode setting provides the system with an initial operating state and serves as a baseline configuration for subsequent optimization and adjustment. By analyzing the characteristics of the first sound source captured in the preset mode, the system can evaluate whether the current preset mode is suitable for a specific meeting environment. For example, if the sound source localization accuracy is insufficient under the preset linear arrangement, the system may need to adjust to a circular arrangement; if the noise suppression effect is not ideal under the preset 8-unit scale, the number of microphone units may need to be increased. This adaptive optimization mechanism, starting with the preset mode, enables the system to find the optimal working configuration based on the characteristics of the actual meeting environment.

[0036] Compared with the traditional fixed-mode microphone array system, the present invention significantly improves the system's adaptability to different conference environments by using a preset mode as the initial configuration and combining it with an adaptive adjustment mechanism.

[0037] Based on the above embodiment, as an optional embodiment, in step 101, the step of obtaining a first sound source collected by a microphone array in a preset mode in the target area may further include the following steps:

[0038] Step 201: Obtain the distribution locations of participants in the target area.

[0039] Specifically, obtaining the distribution of attendees within the target area involves the system determining the spatial distribution of attendees within the target area before meeting check-in begins. This distribution emphasizes the spatial organization and activity characteristics of attendees, rather than simple location coordinates or density. For example, during a meeting, attendees typically maintain a relatively fixed seating arrangement, in which case the distribution can be described as "static." During the check-in phase, attendees may need to queue in a designated area, in which case the distribution is described as "dynamic."

[0040] Furthermore, the purpose of acquiring attendee locations is to identify the application characteristics of the current meeting scene and provide a basis for selecting subsequent sound source acquisition strategies. The system can use visual sensors deployed in the conference room to capture images or videos and analyze the distribution of attendees using image processing algorithms. It can also determine the spatial distribution characteristics of sound sources by analyzing acoustic signals from microphone arrays. Alternatively, it can determine the overall distribution pattern through signal interaction between attendees' smart devices and positioning base stations within the conference room.

[0041] Specifically, the system first needs to establish a scene recognition model. For static distribution scenarios, the system focuses on features such as whether attendees' seating positions are relatively stable and whether attendees are distributed regularly across different areas. For dynamic distribution scenarios, the system needs to identify information such as the location of queues, the direction of queue formation, and the movement patterns of people. This scene feature recognition process is continuous, and the system confirms changes in distribution patterns through real-time monitoring.

[0042] Step 202: When the distribution positions of the participants are statically distributed, the target area is divided into at least one sub-area based on the distribution positions of the participants and the distribution positions of the microphones in the target area, and the first sound source collected by the microphone array in the preset mode in each sub-area is obtained.

[0043] Specifically, if the participants are determined to be statically distributed, the system will divide the target area into sub-areas based on the spatial distribution of the participants and microphones, and then collect sound sources based on this distribution. This static distribution often occurs during the formal meeting phase, when participants are already seated in a relatively fixed position in the conference room, such as around a conference table or in a seating area in a lecture hall.

[0044] Furthermore, the purpose of dividing the target area into sub-areas is to achieve more targeted sound source collection. Since the positions of participants in a static distribution are relatively fixed, the system can divide the target area into several sub-areas with relatively independent acoustic characteristics based on the distribution characteristics of the participants and the deployment positions of the microphones. For example, in a multi-functional conference hall, the target area can be divided into a podium area, a central conference area, and a surrounding audience area; in a large conference room, the area can be divided into several seating areas based on the layout of the conference table. This area division method based on actual scene characteristics can better adapt to the acoustic characteristics of the conference environment.

[0045] Specifically, the system first needs to determine the sub-area division scheme. A Voronoi diagram-based area division algorithm can be used, with the microphone position as the center point, to divide the target area into multiple sub-areas based on the spatial distance relationship. The following factors will be considered during the division: First, the number of participants in each sub-area should be relatively balanced to avoid the situation where there are too many or too few people in a sub-area; second, the boundaries of the sub-areas should be coordinated with the physical layout of the conference room (such as the placement of tables and chairs, the location of aisles, etc.); finally, each sub-area should have sufficient microphone coverage to ensure the quality of sound source collection.

[0046] After completing the area division, the system will use the microphone array in the preset mode to collect sound sources in each sub-area. Since the acoustic environment and the distribution characteristics of participants in each sub-area may be different, the system will adjust the working parameters of the microphone array accordingly. For example, for sub-areas with a denser concentration of participants, it may be necessary to improve the directional performance of the microphone array; while for sub-areas with more dispersed participants, it may be necessary to expand the pickup range of the microphone. The system will continuously monitor the quality of the first sound source collected in each sub-area, including parameters such as signal-to-noise ratio, sound source directionality, and reverberation level, to provide a basis for subsequent mode optimization.

[0047] For example, in a static distribution scenario, it is assumed that the microphone positions in the target area are M1, M2, ..., M N , the positions of the participants are P1, P2, ...P k ,The goal is to divide the target area into several sub-areas so that each microphone is responsible for collecting the sound in its sub-area.

[0048] Furthermore, in order to achieve the best region division effect, the present invention adopts a region division algorithm based on the Voronoi diagram. This algorithm divides each point in space into the region corresponding to the nearest microphone. Specifically, for each microphone M i , and its corresponding Voronoi cell V i Defined by the following conditions:

[0049]

[0050] Among them, d(P,M j ) represents the point P and the microphone M i The Euclidean distance between:

[0051]

[0052] In this way, the system can divide the target area into N non-overlapping sub-areas, and the distance from all points in each sub-area to its corresponding microphone is shorter than the distance to other microphones.

[0053] The Voronoi diagram-based region partitioning method described above ensures that each participant is placed within the nearest microphone collection area, minimizing sound attenuation during propagation. Secondly, because the Voronoi diagram's divisions are based on natural divisions of spatial distance, the resulting sub-region boundaries often adapt well to the physical layout of the conference room. Finally, this partitioning method is computationally simple, produces unique results, and can quickly respond to changes in participant distribution. In this way, the system can establish clear microphone responsibility areas, providing a clear spatial division basis for subsequent sound source collection, thereby improving the efficiency of the entire conference check-in system.

[0054] Step 203: When the distribution positions of the participants are dynamically distributed, the queuing area in the target area is determined based on the distribution positions of the participants and the microphones in the target area, and the first sound source collected by the microphone array in the preset mode in the queuing area is obtained.

[0055] Specifically, if the attendees are dynamically distributed, the system needs to identify the queuing area within the target area based on the spatial distribution of the attendees and microphones, and then collect sound sources in that area. This dynamic distribution often occurs during the initial check-in phase of a meeting, when attendees are not yet seated but are queuing in a specific area, such as at the conference room entrance or waiting in line at the check-in counter.

[0056] Furthermore, the purpose of determining queuing areas is to accurately capture sound source information in dynamic scenarios. Because attendees constantly move while queuing, and the crowds are relatively dense, traditional fixed-area sound source collection methods may not be effective in this dynamic scenario. By identifying and determining queuing areas, the system can focus sound source collection on the spatial range where attendees primarily move around, improving collection efficiency. For example, one or more queues may form at the entrance of a conference room. The system needs to accurately identify the location and direction of these queues and adjust the microphone array's operating mode accordingly.

[0057] Specifically, the system first needs to determine the scope and characteristics of the queuing area through real-time monitoring. This can be done by analyzing attendees' movement trajectories and clustering characteristics using a spatiotemporal clustering approach. This identification process considers the following factors: First, the system analyzes attendees' movement directions to identify clearly directional trajectories; second, it focuses on changes in attendee density, identifying areas with densely populated areas that exhibit linear or strip-like distribution; and finally, the physical layout of the conference room is considered to determine the final scope of the queuing area.

[0058] After determining the queuing area, the system will arrange and adjust the microphone array in the area in a targeted manner. Because the queuing area is characterized by dense crowds and frequent movement, the system uses the microphone array in the preset mode to collect sound sources.

[0059] It should be noted that the coverage of the microphone array needs to match the length and width of the queue; secondly, considering the mobility of the participants, the system may need to enable the sound source tracking function to ensure the continuity of sound collection; finally, in response to the high ambient noise that may occur in the queuing area, the system will appropriately adjust the noise reduction parameters of the microphone array.

[0060] In an optional implementation, when the distribution of participants is dynamically distributed, the system needs to determine the queuing area based on the real-time movement characteristics of the participants. Specifically, assuming that the real-time positions of the participants in the target area are P1(t), P1(t), ..., P k (t), because participants constantly move while queuing, traditional fixed area division methods are difficult to adapt to this dynamic scenario. Therefore, this paper uses a dynamic clustering algorithm based on K-means to determine the queuing area. By tracking and analyzing the changes in participants' positions in real time, it adaptively identifies and updates the scope of the queuing area.

[0061] Furthermore, the system first needs to set the cluster centers C1, C2, ..., C m The value of m can be predetermined based on the conference scale and venue layout. During the algorithm initialization phase, the system randomly selects m participant locations as the initial cluster centers.

[0062] This random selection initialization method can prevent the algorithm from falling into the local optimal solution and improve the global optimality of the clustering results. After that, the system enters the iterative optimization stage: for each participant position P updated in real time, k (t), find the nearest cluster center by calculating the Euclidean distance, that is, solve:

[0063]

[0064] After completing a round of allocation, the system will recalculate the center position of each cluster, and the update formula is:

[0065]

[0066] Among them, S j Belongs to cluster C j The above steps are repeated for all participant locations until the cluster centers converge. The queuing area can be defined by the clustering results.

[0067] This dynamic clustering method can reflect attendees' movement patterns in real time. When attendees' positions change, the cluster centers adjust accordingly, ensuring that the division of queue areas always remains consistent with the actual situation. Secondly, by calculating cluster centers, the system can effectively identify densely populated areas, which are usually where queues occur. Finally, the iterative optimization characteristics of the algorithm ensure that the clustering results will continuously converge to the optimal solution, providing a reliable basis for the accurate division of queue areas.

[0068] In actual applications, the system continuously monitors the changing trends of cluster centers. When the location of the cluster centers stabilizes, the algorithm is considered to have converged, and the clustering results at this point can accurately reflect the spatial distribution of the queuing area. Based on the final clustering results, the system can determine the specific scope of each queuing area and adjust the operating parameters of the microphone array accordingly to ensure the accuracy of voiceprint collection. For example, for identified long queues, the system can deploy a linear microphone array along the queue direction; for more dispersed waiting areas, a ring or matrix microphone layout can be used. This queuing area identification method based on dynamic clustering can effectively improve the system's adaptability and reliability in dynamic scenarios.

[0069] Step 301: Obtain the noise level in the first sound source and the density of participants.

[0070] Among them, the noise level in the first sound source refers to the measurement value of the intensity of all background sounds excluding the effective voice signal in the target area. In the embodiment of the present application, it can be understood as the acoustic interference intensity formed by the superposition of multiple noise sources such as the operating sound of environmental equipment, room reverberation, external interference, etc., which is usually quantified in decibels (dB). For example, the operating noise of the air-conditioning system, the heat dissipation sound of the projection equipment, the traffic noise entering from the outside, and the reverberation effect of these noises in the conference room are all components of the noise level of the first sound source.

[0071] In this embodiment of the present application, the noise level of the first sound source is used to assess the acoustic quality of the conference environment and predict potential interference with voiceprint collection. By measuring and analyzing the noise level, the system can determine whether to adjust the signal processing parameters of the microphone array, such as increasing the signal filtering strength in high-noise environments or reducing the signal gain in low-noise environments to avoid distortion. Furthermore, the changing trend of the noise level can also reflect the dynamic characteristics of the conference environment, providing an important reference for real-time optimization of the microphone array.

[0072] Participant density refers to the number of participants per unit area within the target area. In the present application, this refers to the concentration of participants within the conference space, typically expressed in terms of number of participants per square meter. For example, in a large conference room, the podium area may only have a few speakers, resulting in a low density; whereas the check-in waiting area, where participants gather in queues, may have a higher density.

[0073] In this embodiment of the present application, the density of attendees is used to evaluate the spatial distribution characteristics of sound sources and predict the degree of possible sound source interference. By calculating and analyzing the density of attendees, the system can determine the optimal coverage range and spatial resolution requirements of the microphone array. For example, in areas with high population density, the system needs to improve the spatial selectivity of the microphone array to accurately separate the voices of adjacent attendees; in areas with low population density, the pickup range of a single microphone can be appropriately expanded to improve the system's coverage efficiency.

[0074] For example, in a conference scenario, the acoustic environment in the target area is usually complex, including voice signals of multiple participants and various environmental noises. In order to accurately obtain the voiceprint information of the participants, the system uses a microphone array to collect and process the sound signals. Specifically, when the microphone M i When collecting sound signals in the target area, the acquired signal S i (t) contains the superposition of all sound sources in the area, which can be expressed as:

[0075]

[0076] Where A k is the sound source intensity of the kth participant, ω k is the frequency of the kth sound source, is the phase of the kth sound source, N i (t) is the microphone M i The collected noise.

[0077] In order to improve the quality of sound source acquisition, the system introduces beamforming technology. By weighted superposition of the signals collected by each microphone in the microphone array, the output signal of the beamforming can be expressed as:

[0078]

[0079] Where w i It is a microphone M i The weight, τ i It is a microphone M i The time delay of the acquired signal relative to the reference microphone.

[0080] The system can specifically enhance sound signals from specific directions while suppressing interference from other directions. The weight coefficient determines the contribution of each microphone to the final output signal, while the time delay reflects the difference in the time required for sound to propagate to different microphone positions.

[0081] Among them, the delay τ i The distance d from the sound source to the microphone can be i calculate:

[0082]

[0083] Where c is the speed of sound.

[0084] In practical applications, the system needs to accurately evaluate the acoustic environment characteristics of the target area. First, by analyzing the background noise signal N(t) collected by the microphone array, the noise level L N It can be calculated by the following formula:

[0085]

[0086] The attendee density ρ is defined as the number of attendees per unit area and is calculated as:

[0087]

[0088] Where K is the total number of participants and A is the area of the target area.

[0089] Through this approach, the system can accurately locate sound sources and achieve high-quality audio capture. In areas with high noise levels, the spatial filtering effect can be enhanced by adjusting the weight coefficients; in areas with high participant density, spatial resolution can be improved by optimizing the delay parameters. This adaptive adjustment mechanism based on environmental characteristics improves the system's voiceprint collection performance in complex meeting environments and provides high-quality audio input for subsequent voiceprint recognition.

[0090] Step 302: Adjust the arrangement of the microphone array to a target arrangement and the array size to a target array size based on the noise level and the occupant density.

[0091] The target arrangement and target array size refer to the spatial layout and number of microphone units in the microphone array, which are dynamically adjusted by the system based on the characteristics of the actual meeting environment. In the present application, this can be understood as follows: the target arrangement includes the specific geometric distribution of microphone units in space and the relative spacing between units; the target array size refers to the total number of microphone units participating in the operation after optimization.

[0092] For example, in a specific scenario, the target arrangement may be a matrix arrangement with a spacing of 15 cm and a target array size of 16 microphone units; in another scenario, it may be adjusted to a circular arrangement with a spacing of 30 cm and a target array size of 8 microphone units.

[0093] Furthermore, the target arrangement and target array size are primarily used to achieve optimal adaptation of the microphone array to the conference environment. By adjusting these two parameters, the system can select the most appropriate spatial sampling strategy for different noise levels and crowd density characteristics.

[0094] The target arrangement determines the array's spatial selectivity, affecting the system's ability to distinguish sound sources from different directions. The target array size directly impacts the system's signal processing capabilities, influencing noise suppression and signal enhancement. This dynamic optimization mechanism for parameter combinations ensures the system maintains stable voiceprint collection quality in a variety of meeting scenarios.

[0095] Based on the above embodiment, as an optional embodiment, in step 302, adjusting the arrangement of the microphone array to the target arrangement and the array size to the target array size based on the noise level and the occupant density may further include the following steps:

[0096] Step 401: Determine a target arrangement of a microphone array based on the density of people.

[0097] Specifically, determining the target arrangement of the microphone array based on the density of people refers to the system selecting the optimal spatial arrangement pattern of the microphone units according to the spatial distribution density of the participants in the target area. In the embodiment of the present application, the density ρ of the participants directly reflects the spatial distribution characteristics of the sound source. Under different conditions of personnel density, the system needs to adopt different microphone arrangements to achieve the optimal sound source collection effect. For example, in the check-in queue area with dense participants, the crosstalk problem between adjacent sound sources is more prominent, and a more refined spatial sampling strategy is needed at this time; in areas where participants are dispersed, it is more necessary to consider the uniformity of the sound field coverage.

[0098] Furthermore, the system has designed three typical microphone arrangement patterns, targeting different crowd density scenarios.

[0099] When the occupant density ρ ≥ ρ 1, indicating that the sound sources within the target area are extremely densely distributed, the system adopts a square grid arrangement, with microphone units arranged in an evenly spaced square grid. The spacing between adjacent microphones is typically set to 0.1 to 0.2 meters. This compact orthogonal arrangement provides the maximum spatial sampling density, effectively improving the system's ability to resolve adjacent sound sources.

[0100] When ρ2≤ρ<ρ1, it indicates that the sound source distribution is in a medium-density state. The system selects a triangular grid arrangement mode, and the microphone units are arranged according to a regular triangle grid. This honeycomb arrangement can provide more uniform sound field coverage while maintaining good spatial resolution.

[0101] When ρ<ρ2, it indicates that the sound source distribution is relatively dispersed. The system adopts a linear arrangement mode, and the microphone units are evenly arranged along a straight line. The unit spacing can be appropriately increased to 0.3 to 0.5 meters. This arrangement simplifies the system structure while ensuring the effective collection of dispersed sound sources.

[0102] Step 402: Determine a target array size of the microphone array in a target arrangement based on the noise level.

[0103] After determining the arrangement, the next step is to determine the microphone density. This density refers to the physical spacing between adjacent microphone units in the array, and its value directly impacts the system's spatial sampling accuracy and resolution of sound sources. In conference scenarios, the density of attendees affects the spatial distribution of sound sources, so the system needs to dynamically adjust the microphone density based on the actual crowd density to achieve optimal sound source collection.

[0104] Furthermore, the embodiment of the present application adopts an adaptive arrangement density adjustment method based on personnel density. By establishing a mathematical relationship model between microphone arrangement density and participant density, the system can automatically calculate the most suitable microphone spacing based on the personnel distribution characteristics in the current conference environment. Specifically, the target arrangement density can be expressed as:

[0105]

[0106] Where, d target is the target array density, i.e. the spacing between microphones, d base is the initial arrangement density in the preset mode, ρ0 is the baseline value of the participant density, ρ is the current participant density, and β is the adjustment coefficient, which indicates the degree of influence of the participant density on the arrangement density.

[0107] The above-mentioned dynamic arrangement density adjustment mechanism based on personnel density establishes an inverse relationship between arrangement density and personnel density. The system can automatically improve the spatial sampling accuracy when the participants are dense, and expand the coverage when the participants are sparse, thereby achieving optimal resource utilization; secondly, the introduction of the adjustment coefficient β enables the system to flexibly adjust the response characteristics according to actual application needs. When the β value is large, the system is more sensitive to changes in personnel density. When the β value is small, the arrangement density is maintained relatively stable; finally, this adaptive adjustment mechanism can continuously optimize the spatial layout of the microphone array to ensure that ideal voiceprint collection effects can be obtained in different meeting scenarios.

[0108] After determining the microphone arrangement and density, the next step is to adjust the size of the microphone array based on the noise level. The size of the microphone array refers to the total number of microphone units involved, and its value directly affects the system's signal processing capabilities and noise immunity. In a real-world meeting environment, due to multiple noise sources such as air conditioning systems, equipment operation, and personnel activities, environmental noise can significantly affect the quality of voiceprint signal collection. Therefore, the system needs to dynamically adjust the size of the microphone array based on the current noise level to ensure that the signal-to-noise ratio of the collected signal meets the requirements of voiceprint recognition.

[0109] Furthermore, the embodiments of the present application adopt an adaptive sizing method based on noise levels. By establishing a mathematical relationship model between the number of microphones and ambient noise, the system can automatically calculate the required number of microphones based on the real-time monitored noise level.

[0110] For example, the scale of the target microphone array can be determined using the following formula:

[0111]

[0112] Where N target is the target microphone array size, N base is the initial microphone array size, L0 is the noise threshold preset by the system, and L N is the current noise level, and α is the adjustment coefficient related to the noise level, which indicates the degree of influence of noise on the number of microphones.

[0113] Through a dynamic scale adjustment mechanism based on noise levels, a linear relationship can be established between the number of microphones and the noise level. The system can automatically increase signal acquisition points when the noise is strong and reduce redundant units when the noise is weak, thus realizing on-demand allocation of system resources. Secondly, the introduction of the adjustment coefficient α enables the system to adjust the response strength according to actual application requirements. When the α value is large, the system responds more sensitively to noise changes. When the α value is small, the array size remains relatively stable. Finally, this adaptive adjustment mechanism can continuously optimize the system configuration to ensure sufficient signal quality in different noise environments.

[0114] As an optional embodiment, once the microphone arrangement has been determined, the area coverage of the microphone array can be further determined based on different arrangements. Area coverage refers to the proportion of the microphone array's effective acquisition range within the target area, and its value is directly related to the system's acoustic coverage of the meeting space. In an actual meeting environment, because different microphone arrangements will form different spatial sampling patterns, the total number of microphones required must be calculated based on the specific arrangement to ensure effective coverage of the entire target area.

[0115] The total number of target microphones can be calculated using the following formula:

[0116]

[0117] Where A room is the total area of the target region, A cell is the area of each microphone unit.

[0118] For different arrangements, the unit area covered by the microphone can be calculated using the following formula:

[0119] Square grid arrangement:

[0120] Triangular grid arrangement:

[0121] Linear arrangement: A cell =d target l;

[0122] where l is the length of the microphone coverage area in the linear arrangement.

[0123] The above-mentioned arrangement-based microphone quantity calculation mechanism can establish area coverage models under different arrangement methods. The system can select the optimal arrangement method for different conference space layouts to maximize spatial sampling efficiency; secondly, the introduction of the unit area calculation formula enables the system to accurately evaluate the coverage characteristics of each arrangement method, thereby optimizing the number of microphones used while ensuring the collection effect; finally, this calculation mechanism can be combined with the aforementioned arrangement density adjustment method to achieve overall optimization of microphone array deployment.

[0124] Step 102: When the meeting sign-in begins, obtain a second sound source collected by the microphone array in target mode, obtain a voiceprint recognition result of the second sound source, and perform meeting sign-in based on the recognition result.

[0125] Among them, the second sound source refers to the participant's voice signal collected by the microphone array in target mode during the conference check-in stage. In the embodiment of the present invention, it can be understood as the sound information containing personal voice characteristics emitted by the participant at the designated location. The sound information includes but is not limited to the participant reading out his or her own name, work number or preset password. The second sound source is different from the first sound source collected by the system before the start of check-in to adjust the microphone array. It is mainly used for voiceprint recognition and verification of the participant's identity, and provides voiceprint feature data for subsequent automatic check-in. The sound source has unique acoustic characteristics, including the frequency characteristics, volume, duration, etc. of the sound source. These characteristics together constitute the voiceprint information that can be used for identity recognition.

[0126] Furthermore, the acquisition of the second sound source occurs after the system has optimized the microphone array configuration. At this point, the microphone array has been adjusted to the target mode that best suits the current meeting environment, enabling optimal acquisition of participants' voice signals. By acquiring this second sound source, the system can obtain higher-quality, more distinctive sound signals, thereby improving the accuracy of subsequent voiceprint recognition and ultimately enabling reliable identity verification and automatic sign-in.

[0127] Based on the above embodiment, as an optional embodiment, in step 102, the step of obtaining the voiceprint recognition result of the second sound source may further include the following steps:

[0128] Step 501: Acquire an audio signal corresponding to a second sound source;

[0129] Since the second sound source collected by the microphone array is an analog acoustic signal, it needs to be converted into an audio signal that can be digitally processed for subsequent voiceprint feature extraction and recognition matching. Specifically, the system converts the collected acoustic vibration signal into an electrical signal through each microphone in the microphone array, and then converts the analog electrical signal into a digital audio signal through an analog-to-digital converter.

[0130] During the conversion process, the system uses an appropriate sampling frequency and quantization bit number for audio sampling. Considering that the frequency range of human voice is mainly concentrated between 20Hz and 20kHz, this embodiment selects a sampling frequency of 44.1kHz to ensure that the collected audio signal can fully preserve the acoustic characteristics of human voice. At the same time, using 16-bit quantization accuracy for sampling provides sufficient dynamic range to express subtle changes in the sound, ensuring accurate extraction of voiceprint features.

[0131] This embodiment also incorporates beamforming technology from the microphone array in target mode during audio signal acquisition. By adjusting the beam direction of each microphone, the system can enhance the sound signal from the target direction while suppressing noise interference from other directions. The audio signal acquired in this way has a higher signal-to-noise ratio, which helps improve the accuracy of subsequent voiceprint recognition.

[0132] Step 502: Separate the audio signal based on the time difference between the microphones in the microphone array to obtain a plurality of audio sub-signals.

[0133] Because multiple participants may speak simultaneously in a conference setting, the mixed audio signals collected by the microphone array need to be separated to obtain the independent sound signal of each participant. The system uses the time difference characteristics between the microphones in the microphone array to separate the audio signals. This time difference-based separation method can effectively solve the problem of sound source aliasing.

[0134] Specifically, assuming that the signal emitted by the sound source is S(t), the distances between the participants and the microphones are different, resulting in differences in the sound propagation time. i The collected signal S i (t) is actually the result of the original signal after time delay, that is:

[0135] S i (t) = S(t-τ i );

[0136] In order to accurately separate the independent audio signals, this embodiment uses the delay and summation method for signal processing. First, the system performs delay compensation on the signal collected by each microphone and adjusts it to:

[0137] S′ i (t) = S i (t+τ i ) = S(t);

[0138] This allows the signals from all microphones to be time-synchronized. This delay compensation ensures that signals from the same sound source can be coherently added, while signals from different sources will be weakened due to time differences.

[0139] After completing the delay compensation, the system performs weighted superposition processing on the synchronized signals to obtain the enhanced signal:

[0140]

[0141] where w i It is a microphone M i The weights are determined based on the performance characteristics of each microphone, primarily considering factors such as the microphone's signal-to-noise ratio and frequency response. Microphones with higher signal-to-noise ratios are given greater weight to improve the quality of the separated signal.

[0142] In a lab test environment, when three participants spoke simultaneously, the system was able to separate the mixed signal into three independent audio sub-signals, improving the signal-to-noise ratio by approximately 8dB. Even in a noisy environment, the separated audio sub-signals maintained high quality, with clearly discernible voiceprint features. This separation provides a reliable signal foundation for subsequent voiceprint recognition, significantly improving recognition accuracy when multiple participants sign in simultaneously.

[0143] This time-difference-based signal separation method not only solves the aliasing problem encountered by traditional conference sign-in systems when multiple people sign in simultaneously, but also improves the real-time performance of the system's audio signal processing. The processing delay of the entire signal separation process is kept to less than 100 milliseconds, meeting the real-time requirements of conference sign-in systems. Furthermore, this method has relatively low computational complexity, making it suitable for implementation in embedded systems and possessing excellent engineering practicality.

[0144] Step 503: Match the voiceprint recognition results corresponding to each audio sub-signal.

[0145] To achieve accurate identity recognition, the system needs to extract and match the voiceprint features of each separated audio sub-signal. The core of voiceprint recognition is to convert the audio signal into a feature vector that can uniquely represent the speaker's identity characteristics and match it with pre-stored voiceprint features.

[0146] Optionally, this embodiment uses the MFCC method to extract voiceprint features. First, the audio sub-signal is framed, and the length of each frame is set to 25ms, and the frame shift is 10ms to ensure the continuity of signal analysis and the local characteristics of the time domain. After each frame signal is processed with a Hamming window, the spectrum is obtained by fast Fourier transform (FFT). Taking into account the human ear's perception characteristics of sounds of different frequencies, the system converts the spectrum to a Mel frequency scale that is more in line with the human ear's auditory characteristics. Subsequently, the Mel spectrum is logarithmically processed to simulate the human ear's nonlinear perception characteristics of sound intensity. Finally, the logarithmic Mel spectrum is converted into an MFCC feature vector V by discrete cosine transform (DCT).i , the feature vector contains the speaker's vocal tract feature information.

[0147] When performing feature matching, the system adopts a dual judgment mechanism, that is, it uses both Euclidean distance and cosine similarity to calculate the similarity of feature vectors. i , the system calculates the feature vector V of the M participants registered in the voiceprint database j Euclidean distance and cosine similarity between (j=1,2,…,M).

[0148] Among them, the Euclidean distance expression is:

[0149]

[0150] Where d is the dimension of the feature vector, V i,k and V j,k They are the eigenvectors V i and V j The kth component of .

[0151] The cosine similarity expression is:

[0152]

[0153] Where V i ·V j is the dot product of two vectors, ||V i ||||V j || is V i ·V j The norm of .

[0154] Furthermore, this embodiment sets a strict double-threshold decision rule: only when the Euclidean distance is less than the preset threshold θ d And the cosine similarity is greater than the preset threshold θ s A match is considered successful only when the feature vector matches the speaker's identity. This dual-judgment mechanism effectively reduces false positives and improves system reliability. When a matching feature vector is found, the system confirms the speaker's identity and automatically completes the sign-in record.

[0155] Reference Figure 2 , this application also provides a conference check-in system based on voiceprint recognition, including:

[0156] A microphone array adjustment module is configured to obtain a first sound source collected by a microphone array in a preset mode in a target area when the conference check-in has not yet begun, and adjust the preset mode to a target mode based on the first sound source, wherein the preset mode includes an arrangement and array size of the microphone array;

[0157] The voiceprint recognition sign-in module is used to obtain a second sound source collected by the microphone array in the target mode when the meeting sign-in starts, obtain the voiceprint recognition result of the second sound source, and perform meeting sign-in based on the recognition result.

[0158] On the basis of the above embodiment, as an optional embodiment, the microphone array adjustment module is further used to obtain the distribution positions of the participants in the target area; when the distribution positions of the participants are statically distributed, based on the distribution positions of the participants and the distribution positions of the microphones in the target area, the target area is divided into at least one sub-area, and the first sound source collected by the microphone array in a preset mode in each sub-area is obtained; when the distribution positions of the participants are dynamically distributed, based on the distribution positions of the participants and the microphones in the target area, the queuing area in the target area is determined, and the first sound source collected by the microphone array in a preset mode in the queuing area is obtained.

[0159] Based on the above embodiments, as an optional embodiment, the microphone array adjustment module is further used to obtain the noise level in the first sound source and the density of participants; based on the noise level and the density of participants, the arrangement of the microphone array is adjusted to a target arrangement and the array size is adjusted to a target array size; the target arrangement and the target array size are determined as a target mode, and the preset mode is adjusted to the target mode; wherein, the noise level in the third sound source collected by the microphone array in the target mode is less than the corresponding threshold, and the time difference between the microphones in the microphone array is greater than the corresponding threshold.

[0160] Based on the above embodiment, as an optional embodiment, the microphone array adjustment module is further used to determine the target arrangement of the microphone array based on the personnel density; and determine the target array size of the microphone array under the target arrangement based on the noise level.

[0161] Based on the above embodiment, as an optional embodiment, the voiceprint recognition sign-in module is further used to obtain the audio signal corresponding to the second sound source; separate the audio signal based on the time difference between each microphone in the microphone array to obtain multiple audio sub-signals; and match the voiceprint recognition results corresponding to each of the audio sub-signals.

[0162] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0163] An embodiment of the present application also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded by a processor and executed by a conference check-in method based on voiceprint recognition as described in the above embodiment. The specific execution process can refer to the specific description of the embodiment shown and will not be repeated here.

[0164] This application also discloses an electronic device. Figure 3 , Figure 3 The electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .

[0165] The communication bus 302 is used to implement the connection and communication between these components.

[0166] The user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0167] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0168] The processor 301 may include one or more processing cores. The processor 301 utilizes various interfaces and circuits to connect various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one hardware form selected from the group consisting of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface graphics, and applications; the GPU is responsible for rendering and drawing the content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 and may be implemented separately on a single chip.

[0169] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may optionally be at least one storage device located away from the aforementioned processor 301. Refer to Figure 3 , as a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface module, and an application program for a conference check-in method based on voiceprint recognition.

[0170] exist Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 305 for a conference check-in method based on voiceprint recognition. When executed by one or more processors 301, the electronic device 300 executes one or more of the methods described in the above embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.

[0171] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the conference check-in method based on voiceprint recognition provided by the above methods.

[0172] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0173] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0174] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0175] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0176] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.

[0177] The foregoing is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.

[0178] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.

Claims

1. A conference sign-in method based on voiceprint recognition, characterized in that: include: In the case that the conference check-in has not started, the distribution positions of the participants in the target area are obtained; in the case that the distribution positions of the participants are statically distributed, based on the distribution positions of the participants and the distribution positions of the microphones in the target area, the target area is divided into at least one sub-area, and the first sound source collected by the microphone array in the preset mode in each sub-area is obtained; the noise level in the first sound source and the density of the participants are obtained; based on the noise level and the density of the participants, the arrangement mode of the microphone array is adjusted to the target arrangement mode and the array size is adjusted to the target array size; the target arrangement mode and the target array size are determined to be the target mode, and the preset mode is adjusted to the target mode; wherein the noise level in the third sound source collected by the microphone array in the target mode is less than the corresponding threshold, and the time difference between the microphones in the microphone array is greater than the corresponding threshold; the preset mode includes the arrangement mode and the array size of the microphone array; When the conference sign-in starts, a second sound source collected by the microphone array in the target mode is obtained, a voiceprint recognition result of the second sound source is obtained, and the conference sign-in is performed based on the voiceprint recognition result.

2. The conference sign-in method based on voiceprint recognition according to claim 1 is characterized in that: The method further comprises: In the case where the distribution positions of the participants are dynamically distributed, a queuing area in the target area is determined based on the distribution positions of the participants and microphones in the target area, and a first sound source collected by the microphone array in a preset mode in the queuing area is obtained.

3. The conference check-in method based on voiceprint recognition according to claim 1 is characterized in that: The step of adjusting the arrangement of the microphone array to a target arrangement and the array size to a target array size based on the noise level and the personnel density includes: determining a target arrangement of the microphone array based on the density of people; A target array size of the microphone array in the target arrangement is determined based on the noise level.

4. The conference sign-in method based on voiceprint recognition according to claim 1 is characterized in that: After obtaining the second sound source collected by the microphone array in the target mode, the method further includes: Sound field information corresponding to the second sound source is constructed, and beam directions of microphones in the microphone array in the target mode are adjusted based on the sound field information.

5. The conference sign-in method based on voiceprint recognition according to any one of claims 1 to 4, characterized in that: The obtaining of the voiceprint recognition result of the second sound source includes: Obtaining an audio signal corresponding to the second sound source; Separating the audio signal based on the time difference between the microphones in the microphone array to obtain a plurality of audio sub-signals; Match the voiceprint recognition results corresponding to each of the audio sub-signals.

6. A conference check-in system based on voiceprint recognition, characterized in that: include: A microphone array adjustment module is used to obtain the distribution positions of participants in a target area when the conference check-in has not started; when the distribution positions of the participants are statically distributed, based on the distribution positions of the participants and the distribution positions of the microphones in the target area, divide the target area into at least one sub-area, and obtain a first sound source collected by the microphone array in a preset mode in each sub-area; obtain the noise level and the density of participants in the first sound source; adjust the arrangement mode of the microphone array to a target arrangement mode and the array size to a target array size based on the noise level and the density of participants; determine the target arrangement mode and the target array size as a target mode, and adjust the preset mode to the target mode; wherein the noise level in the third sound source collected by the microphone array in the target mode is less than the corresponding threshold, and the time difference between the microphones in the microphone array is greater than the corresponding threshold; the preset mode includes the arrangement mode and array size of the microphone array; The voiceprint recognition sign-in module is used to obtain a second sound source collected by the microphone array in the target mode when the meeting sign-in starts, obtain the voiceprint recognition result of the second sound source, and perform meeting sign-in based on the voiceprint recognition result.

7. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 5.

8. A computer storage medium, characterized in that The computer storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 5 is executed.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Metering calibration method and system for sound source identification positioning system

    CN118731848A