A pickup and amplification method and system based on human body recognition and voiceprint matching

By using human body recognition and voiceprint matching technology, the parameters of the sound pickup device are adjusted in real time, which solves the problem of sound interference in multi-person speaking scenarios, ensures clear audio for the speaker, and improves sound pickup and amplification efficiency and user experience.

CN115866499BActive Publication Date: 2026-06-02GUANGZHOU BAOLUN ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU BAOLUN ELECTRONICS CO LTD
Filing Date
2022-12-02
Publication Date
2026-06-02

Smart Images

  • Figure CN115866499B_ABST
    Figure CN115866499B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on human body recognition and voiceprint matching pickup sound amplification method and system, comprising: the human body image of real-time acquisition is carried out human body recognition, obtains the identity ID corresponding to all on-site personnel in current scene;When any identity ID meets ID information base, control is opened and the parameter of pickup sound equipment is adjusted, to collect first audio data, and obtain the voiceprint matching score of the multiple second audio data corresponding to first audio data and main speaker voice;Based on the size relation of each voiceprint matching score and preset threshold value, each second audio data is respectively enhanced or inhibited processing, and the third audio data obtained after processing is amplified.The application is based on the parameter of pickup sound equipment that human body recognition result based on human body image controls adjustment to improve pickup sound effect, and according to the matching degree of first audio data and main speaker prestorage voice, each second audio data is enhanced or inhibited processing, to improve the amplification intelligibility of main speaker voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent sound pickup and amplification technology, and in particular to a sound pickup and amplification method and system based on human body recognition and voiceprint matching. Background Technology

[0002] In daily activities, it's common to encounter situations where multiple people speak simultaneously or engage in continuous dialogue. For example, in a classroom, a teacher might ask a question, and the student would answer; or students might discuss a problem together. For the former, in online classes, the current solution is usually that after the teacher calls on a student, the student being called on manually turns on their microphone, or the teacher manually turns on the student's microphone. To avoid interference between speakers, preventing other students from hearing clearly, the microphone of the previous speaker must be muted when switching speakers. This process is cumbersome, prone to errors, and doesn't address the issue of multiple speakers speaking simultaneously or the hierarchy of speakers. In offline classes, the current solution is usually for the teacher to have a portable or handheld microphone, which is then passed to the student being asked a question. This is both time-consuming and laborious. Even with multiple handheld microphones, interference from multiple speakers speaking simultaneously still occurs, and the issue of hierarchy of speakers remains unresolved, especially in noisy environments, outdoors, or with multiple speakers. Summary of the Invention

[0003] This invention provides a sound pickup and amplification method and system based on human body recognition and voiceprint matching. In scenarios with multiple speakers, it optimizes the sound pickup performance of the pickup device and avoids the impact of non-speaker voices and ambient noise on the clarity of the speaker's voice amplification.

[0004] To address the aforementioned technical problems, embodiments of the present invention provide a sound pickup and amplification method based on human body recognition and voiceprint matching, comprising:

[0005] The system acquires real-time images of the human body to be identified and the location of the sound source, and performs human body recognition on the human body image to obtain the identity IDs of all on-site personnel in the current scene.

[0006] When any of the identity IDs matches the preset ID information database, the parameters of the sound pickup device in the current scene are controlled to be turned on and adjusted so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time, and perform voiceprint recognition and matching on the first audio data according to the second voiceprint features to obtain multiple second audio data and the voiceprint matching score corresponding to each second audio data.

[0007] The second audio data corresponding to the voiceprint matching score that is greater than a preset threshold is enhanced, and the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold is suppressed, so as to obtain the third audio data corresponding to the first audio data, and the third audio data is amplified.

[0008] The ID information database includes user IDs corresponding to all users with speaking privileges, and the second voiceprint feature is obtained by voiceprint recognition of the speaker's pre-stored voice, where the speaker is the user with the highest speaking privileges.

[0009] Implementing this invention involves performing human recognition on real-time acquired human images to determine the identity IDs of all individuals in the current scene. When an identity ID matches a preset ID database, the parameters of the audio pickup device are activated and adjusted. This ensures that the audio pickup device only collects the first audio data corresponding to the sound source location when a user with speaking authority is present, preventing invalid audio data from occupying system storage space and affecting overall audio pickup and amplification efficiency. Furthermore, automatically adjusting the audio pickup device parameters based on the sound source location in the current scene effectively improves the audio pickup and amplification effect, enhancing the overall user experience. In addition, after collecting the first audio data, voiceprint recognition and matching are performed on the first audio data based on the second voiceprint characteristics corresponding to the user with the highest speaking authority. Audio data with voiceprint matching scores higher than a preset threshold is enhanced, while other audio data is suppressed. This achieves gain on the speaker's audio data and filtering of other audio data, ensuring that the user hears the speaker's audio data clearly and preventing interference from other people's audio data.

[0010] As a preferred embodiment, the step of enhancing the second audio data corresponding to the voiceprint matching score greater than a preset threshold and suppressing the second audio data corresponding to the voiceprint matching score less than or equal to the preset threshold to obtain the third audio data corresponding to the first audio data, and then amplifying the third audio data, specifically involves:

[0011] Determine whether the voiceprint matching score of each voiceprint is greater than a preset threshold;

[0012] If so, then the second audio data corresponding to the current voiceprint matching score is enhanced.

[0013] If not, then the second audio data corresponding to the current voiceprint matching score is suppressed;

[0014] All the second audio data that have undergone enhancement or suppression processing are subjected to audio order preservation processing to obtain the fourth audio data corresponding to the first audio data.

[0015] The fourth audio data is subjected to time delay correction processing by a synchronization device to obtain the third audio data corresponding to the first audio data, and the third audio data is amplified; wherein, the synchronization device includes a time delayer, a filter and a generator.

[0016] In a preferred embodiment of the present invention, all second audio data that have undergone enhancement or suppression processing are subjected to audio order preservation processing to ensure the preservation and gain of the speaker's audio data. Then, through a synchronization device including a delay unit, a filter, and a generator, the fourth audio data is subjected to delay correction processing to achieve amplification synchronization and prevent the problem of unclear amplification caused by delay.

[0017] As a preferred embodiment, the sound pickup and amplification method based on human body recognition and voiceprint matching further includes:

[0018] Real-time acquisition of the first action information corresponding to the speaker;

[0019] When the first action information matches a preset first action signal or the first audio data matches a preset voice signal, the second action information corresponding to all the on-site personnel is obtained in real time.

[0020] When the second action information matches the preset second action signal, the identity ID of the on-site personnel corresponding to the current second action information is displayed on the display screen in real time. Then, when the right to speak is granted by the display screen, the parameters of the sound pickup device are controlled and adjusted according to the right to speak signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time and amplify the fifth audio data.

[0021] In a preferred embodiment of the present invention, when the speaker with the highest speaking authority is present, the speaker can enter a speaking authority granting mode by performing an action matching a preset first action signal or speaking a voice matching a preset voice signal. In this mode, if all on-site personnel in the current scene are detected to perform an action matching a preset second action signal, the identity ID of the on-site personnel performing the above action is displayed on the screen in real time. Then, upon receiving the speaking authority granting signal sent by the screen, the parameters of the sound pickup device in the current scene are adaptively adjusted so that the audio data of the on-site personnel performing the above action is picked up and amplified, thereby realizing the speaker's granting of speaking authority to the on-site personnel in the current scene.

[0022] As a preferred embodiment, the step of performing voiceprint recognition and matching on the first audio data based on the second voiceprint features to obtain multiple second audio data and a voiceprint matching score corresponding to each second audio data specifically involves:

[0023] Feature extraction is performed on the first audio data to obtain the audio features corresponding to the first audio data. Based on the audio features, the first audio data is segmented and clustered to obtain multiple sets of first voiceprint features and the second audio data corresponding to each first voiceprint feature.

[0024] Each of the first voiceprint features and the second voiceprint features is matched to obtain the voiceprint matching score corresponding to each of the first voiceprint features, which is then used as the voiceprint matching score corresponding to each of the second audio data.

[0025] In a preferred embodiment of the present invention, the first audio data is sequentially subjected to feature extraction, segmentation and clustering processing. The first audio data is divided into multiple second audio data according to different first voiceprint features. Then, each first voiceprint feature is matched with the second voiceprint features of the speaker to determine the voiceprint matching score corresponding to each second audio data, so as to find the audio data with a high degree of matching with the speaker's voice.

[0026] As a preferred embodiment, the real-time acquisition of the human image and sound source location to be identified, and the performance of human recognition on the human image to obtain the identity IDs of all personnel in the current scene, specifically involves:

[0027] Using a human body tracking algorithm from a camera, the image of the human body to be identified is captured in real time, and the location of the sound source is obtained in real time.

[0028] The human image is subjected to posture recognition, and the recognized posture recognition result is matched with the pre-stored user posture information to obtain the identity IDs corresponding to all the on-site personnel in the current scene.

[0029] A preferred embodiment of the present invention utilizes a human body tracking algorithm to achieve real-time capture of human body images and real-time acquisition of sound source locations. Based on the matching between the human body posture recognition results corresponding to the human body images and the pre-stored user posture information, the identity IDs of all on-site personnel in the current scene are determined, thereby knowing whether there are people with speaking permissions in the current scene, so as to control and adjust the state of the sound pickup device.

[0030] As a preferred embodiment, the sound pickup and amplification method based on human body recognition and voiceprint matching further includes:

[0031] When all the voiceprint matching scores are less than or equal to the preset threshold, a prompt signal is sent to the display screen to remind the speaker to speak.

[0032] In a preferred embodiment of the present invention, when all voiceprint matching scores are less than or equal to a preset threshold, it indicates that the sound pickup device has not collected the speaker's audio data. At this time, it is possible that the speaker has not spoken. Therefore, by sending a prompt signal to the display screen, the speaker is reminded to speak, thereby indirectly advancing the process of the meeting, class, etc.

[0033] To address the same technical problem, embodiments of the present invention also provide a sound pickup and amplification system based on human body recognition and voiceprint matching, comprising:

[0034] The data acquisition module is used to acquire the human image and sound source location to be identified in real time, and to perform human recognition on the human image to obtain the identity IDs of all on-site personnel in the current scene.

[0035] The identification and matching module is used to control the activation and adjustment of the parameters of the sound pickup device in the current scene when any of the identity IDs matches the preset ID information database. This enables the sound pickup device to collect first audio data corresponding to the sound source location in real time, and to perform voiceprint recognition and matching on the first audio data based on the second voiceprint features, thereby obtaining multiple second audio data and a voiceprint matching score corresponding to each second audio data. The ID information database includes user IDs corresponding to all users with speaking permissions, and the second voiceprint features are obtained by voiceprint recognition of the speaker's pre-stored voice. The speaker is the user with the highest speaking permissions.

[0036] The first amplification module is used to enhance the second audio data corresponding to the voiceprint matching score that is greater than a preset threshold, and to suppress the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold, so as to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data.

[0037] As a preferred embodiment, the first amplification module specifically includes:

[0038] An audio adjustment unit is used to determine whether each of the voiceprint matching scores is greater than a preset threshold; if so, it enhances the second audio data corresponding to the current voiceprint matching score; if not, it suppresses the second audio data corresponding to the current voiceprint matching score.

[0039] An audio order preservation unit is used to perform audio order preservation processing on all the second audio data that has undergone enhancement processing or suppression processing, so as to obtain the fourth audio data corresponding to the first audio data.

[0040] A delay correction unit is used to perform delay correction processing on the fourth audio data through a synchronization device to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data; wherein, the synchronization device includes a delay unit, a filter, and a generator.

[0041] As a preferred embodiment, the sound pickup and amplification system based on human body recognition and voiceprint matching further includes:

[0042] The reminder module is used to send a reminder signal to the display screen when all the voiceprint matching scores are less than or equal to the preset threshold, so as to remind the speaker to speak.

[0043] The second amplification module is used to acquire the first action information corresponding to the speaker in real time; when the first action information matches the preset first action signal or the first audio data matches the preset voice signal, it acquires the second action information corresponding to all the on-site personnel in real time; when the second action information matches the preset second action signal, it displays the identity ID corresponding to the on-site personnel corresponding to the current second action information on the display screen in real time, and then, upon receiving the speaking right granting signal sent by the display screen, it controls and adjusts the parameters of the sound pickup device according to the speaking right granting signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time, and amplify the fifth audio data.

[0044] As a preferred embodiment, the identification and matching module specifically includes:

[0045] The device control unit is used to control the activation and adjustment of the parameters of the sound pickup device in the current scene when any of the identity IDs matches the ID information database, so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time;

[0046] The feature extraction unit is used to extract features from the first audio data, obtain audio features corresponding to the first audio data, and perform segmentation and clustering processing on the first audio data according to the audio features to obtain multiple sets of first voiceprint features and second audio data corresponding to each first voiceprint feature.

[0047] The voiceprint matching unit is used to match each of the first voiceprint features with the second voiceprint features to obtain the voiceprint matching score corresponding to each of the first voiceprint features, which is used as the voiceprint matching score corresponding to each of the second audio data. Attached Figure Description

[0048] Figure 1This is a flowchart illustrating a sound pickup and amplification method based on human body recognition and voiceprint matching provided in Embodiment 1 of the present invention.

[0049] Figure 2 : This is a schematic diagram of a sound pickup and amplification system based on human body recognition and voiceprint matching provided in Embodiment 1 of the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Example 1:

[0052] Please refer to Figure 1 This invention provides a sound pickup and amplification method based on human body recognition and voiceprint matching, which includes steps S1 to S3, each of which is detailed below:

[0053] Step S1: Real-time acquisition of the human image and sound source location to be identified, and human recognition of the human image to obtain the identity IDs of all on-site personnel in the current scene.

[0054] As a preferred embodiment, step S1 includes steps S11 to S12, each of which is detailed below:

[0055] Step S11: Using the camera's human body tracking algorithm, capture the image of the human body to be identified in real time and obtain the location of the sound source in real time.

[0056] Step S12: Perform posture recognition on the human image and match the obtained posture recognition results with the pre-stored user posture information to obtain the identity IDs of all on-site personnel in the current scene.

[0057] In this embodiment, before executing step S1, all users register and log in to the system, and their pre-stored voice and body information are pre-entered into the system's database, thereby generating user voiceprint features, user posture information, and user IDs for each user. Next, the administrator sets the speaking permissions and their levels for each user in the system backend to determine the ID information database and the speaker.

[0058] In addition, speaking permissions can be enabled not only through administrator settings in the system backend, but also through user-defined scenarios on the system, such as meeting room reservations. Once a reservation is successful, the system will grant the user and other users who need to speak at the meeting the corresponding speaking permissions based on the reservation information.

[0059] It should be noted that the current scenario includes, but is not limited to, classrooms, meetings, speeches, and skit performances, and the current scenario mode can be online or offline. By accessing the scenario management system, the corresponding scenario mode can be obtained.

[0060] Step S2: When any identity ID matches the preset ID information database, control the activation and adjust the parameters of the sound pickup device in the current scene so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time, and perform voiceprint recognition and matching on the first audio data according to the second voiceprint feature to obtain multiple second audio data and the voiceprint matching score corresponding to each second audio data; wherein, the ID information database includes the user IDs corresponding to all users with speaking permissions, and the second voiceprint feature is obtained by voiceprint recognition of the speaker's pre-stored voice, and the speaker is the user with the highest speaking permission.

[0061] As a preferred embodiment, step S2 includes steps S21 to S23, each of which is detailed below:

[0062] Step S21: When any identity ID matches the ID information database, control to turn on and adjust the parameters of the sound pickup device in the current scene so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time.

[0063] In this embodiment, the sound pickup device includes multiple microphones. If the current scenario is online, the microphone corresponding to the sound source location is activated, and its pickup sensitivity is adjusted. If the current scenario is offline, the microphone is activated according to its type, and its parameters are adaptively adjusted. Specifically, when directional pickup is used, the angle of the pickup device is adjusted so that the pickup device points to the sound source location. When an omnidirectional microphone is used, the sensitivity of the pickup device is adjusted according to the distance between the pickup device and the sound source location.

[0064] It's important to note that in offline scenarios, microphones are typically suspended to free the speaker's hands and eliminate the need for them to wear amplification equipment. The specific process for setting up these microphones in offline scenarios can be abstracted as a graph composed of discrete squares. Each square represents the microphone's reliability after its placement, with reliability scored out of 100. Then, an intelligent algorithm adjusts the number and arrangement of microphones in the offline scenario to ensure that the reliability of each square exceeds a preset value. Finally, the scheme with the fewest microphones and the simplest placement is selected as the final microphone setup for the offline scenario.

[0065] Step S22: Extract features from the first audio data to obtain the audio features corresponding to the first audio data, and perform segmentation and clustering processing on the first audio data according to the audio features to obtain multiple sets of first voiceprint features and second audio data corresponding to each first voiceprint feature.

[0066] Step S23: Match each first voiceprint feature with the second voiceprint feature to obtain the voiceprint matching score corresponding to each first voiceprint feature, which is used as the voiceprint matching score corresponding to each second audio data.

[0067] Step S3: Enhance the second audio data corresponding to the voiceprint matching score that is greater than the preset threshold, and suppress the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold, so as to obtain the third audio data corresponding to the first audio data, and amplify the third audio data.

[0068] As a preferred embodiment, step S3 includes steps S31 to S35, each of which is detailed below:

[0069] Step S31: Determine whether each voiceprint matching score is greater than a preset threshold; if yes, proceed to step S32; otherwise, proceed to step S33.

[0070] Step S32: Enhance the second audio data corresponding to the current voiceprint matching score.

[0071] Step S33: Suppress the second audio data corresponding to the current voiceprint matching score.

[0072] Step S34: Perform audio order preservation processing on all the second audio data that have undergone enhancement or suppression processing to obtain the fourth audio data corresponding to the first audio data.

[0073] In this embodiment, when the second audio data completes enhancement or suppression processing, it enters the input area of ​​the order-preserving module. Then, through the buffer and output area, it forms the order-preserving fourth audio data. At this point, if the second audio data number 1 is about to enter the buffer from the input area, it is directly output to the output area. Next, if the second audio data number 4 enters the buffer, it waits on track 1. Then, if the second audio data number 3 enters the buffer, the buffer automatically allocates a new memory space—track 2—for it, allowing it to wait on track 2. Then, if the second audio data number 2 enters the buffer, it is directly output to the output area, and the second audio data numbers 3 and 4 from the buffer are output to the output area sequentially. This process continues until the fourth audio data corresponding to the first audio data is finally generated in the output area.

[0074] Step S35: The fourth audio data is subjected to time delay correction processing through a synchronization device to obtain the third audio data corresponding to the first audio data, and the third audio data is amplified; wherein, the synchronization device includes a time delayer, a filter and a generator.

[0075] In this embodiment, the delay unit includes a low-speed channel and a high-speed channel. The low-speed channel is connected to the generator, and the high-speed channel is connected to the filter. The low-speed channel optimizes and applies effects to the fourth audio data. Then, the generator optimizes the fourth audio data based on the output of the low-speed channel and inputs the optimized audio data to the filter. The filter removes the additional tags from the aforementioned processing and combines this with the gain audio data from the high-speed channel to obtain the third audio data corresponding to the first audio data.

[0076] As a preferred embodiment, the sound pickup and amplification method based on human body recognition and voiceprint matching provided by the present invention further includes step S4, as follows:

[0077] Step S4: When all voiceprint matching scores are less than or equal to the preset threshold, it indicates that the current sound pickup device has not collected the speaker's audio data. Therefore, a prompt signal is sent to the display screen to remind the speaker to speak, thereby advancing the progress of the meeting, class, etc.

[0078] As a preferred embodiment, the sound pickup and amplification method based on human body recognition and voiceprint matching provided by the present invention further includes steps S5 to S7, the specific steps of which are as follows:

[0079] Step S5: Obtain the first action information corresponding to the speaker in real time.

[0080] Step S6: When the first action information matches the preset first action signal or the first audio data matches the preset voice signal, the second action information corresponding to all on-site personnel is obtained in real time.

[0081] Step S7: When the second action information matches the preset second action signal, the identity ID of the on-site personnel corresponding to the current second action information is displayed on the display screen in real time. Then, when the speaking right granting signal is received from the display screen, the parameters of the sound pickup device are controlled and adjusted according to the speaking right granting signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time and amplify the fifth audio data.

[0082] In this embodiment, when the speaker's first action information matches a preset first action signal or the first audio data matches a preset voice signal, it indicates that the speaker will grant speaking rights. The system then enters the speaking rights granting mode. In this mode, if the second action information of any person in the current scene matches a preset second action signal, it indicates that the person requests to speak, and their corresponding ID is displayed on the screen in real time. The speaker then views all the IDs of those requesting to speak on the screen and selects one to speak by clicking on the screen, causing the screen to send a speaking rights granting signal to the system. Then, based on the received speaking rights granting signal, the system controls and adjusts parameters such as the orientation and sensitivity of the microphone to allow the microphone to collect the audio data of the person granted speaking rights in real time.

[0083] Please refer to Figure 2 This is a schematic diagram of a sound pickup and amplification system based on human body recognition and voiceprint matching, provided by an embodiment of the present invention. The system includes a data acquisition module M1, a recognition and matching module M2, and a first amplification module M3, the specific details of which are as follows:

[0084] The data acquisition module M1 is used to acquire the human images and sound source locations to be identified in real time, and to perform human recognition on the human images to obtain the identity IDs of all on-site personnel in the current scene.

[0085] The identification and matching module M2 is used to control the activation and adjustment of the parameters of the sound pickup device in the current scene when any identity ID matches the preset ID information database. This enables the sound pickup device to collect the first audio data corresponding to the sound source location in real time, and to perform voiceprint recognition and matching on the first audio data based on the second voiceprint features to obtain multiple second audio data and the voiceprint matching score corresponding to each second audio data. The ID information database includes the user IDs corresponding to all users with speaking permissions, and the second voiceprint features are obtained by voiceprint recognition of the speaker's pre-stored voice. The speaker is the user with the highest speaking permissions.

[0086] The first amplification module M3 is used to enhance the second audio data corresponding to the voiceprint matching score that is greater than a preset threshold, and to suppress the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold, so as to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data.

[0087] As a preferred embodiment, the first amplification module M3 specifically includes an audio adjustment unit 31, an audio sequence preservation unit 32, and a time delay correction unit 33, the details of which are as follows:

[0088] The audio adjustment unit 31 is used to determine whether each voiceprint matching score is greater than a preset threshold; if so, it enhances the second audio data corresponding to the current voiceprint matching score; if not, it suppresses the second audio data corresponding to the current voiceprint matching score.

[0089] The audio order preservation unit 32 is used to perform audio order preservation processing on all second audio data that have undergone enhancement processing or suppression processing, so as to obtain the fourth audio data corresponding to the first audio data.

[0090] The delay correction unit 33 is used to perform delay correction processing on the fourth audio data through a synchronization device to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data; wherein, the synchronization device includes a delay unit, a filter and a generator.

[0091] For the preferred option, please refer to Figure 2 The present invention provides a sound pickup and amplification system based on human body recognition and voiceprint matching, which further includes a reminder module M4 and a second amplification module M5. The specific details of each module are as follows:

[0092] The reminder module M4 is used to send a reminder signal to the display screen when all voiceprint matching scores are less than or equal to a preset threshold, so as to remind the speaker to speak.

[0093] The second amplification module M5 is used to acquire the first action information corresponding to the speaker in real time; when the first action information matches the preset first action signal or the first audio data matches the preset voice signal, it acquires the second action information corresponding to all on-site personnel in real time; when the second action information matches the preset second action signal, it displays the identity ID of the on-site personnel corresponding to the current second action information on the display screen in real time, and then, upon receiving the speaking right granting signal sent by the display screen, it controls and adjusts the parameters of the sound pickup device according to the speaking right granting signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time, and amplify the fifth audio data.

[0094] As a preferred embodiment, the identification and matching module M2 specifically includes a device control unit 21, a feature extraction unit 22, and a voiceprint matching unit 23, the details of which are as follows:

[0095] The device control unit 21 is used to control the activation and adjust the parameters of the sound pickup device in the current scene when any identity ID matches the ID information database, so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time.

[0096] The feature extraction unit 22 is used to extract features from the first audio data, obtain the audio features corresponding to the first audio data, and perform segmentation and clustering processing on the first audio data according to the audio features to obtain multiple sets of first voiceprint features and second audio data corresponding to each first voiceprint feature.

[0097] The voiceprint matching unit 23 is used to match each first voiceprint feature with the second voiceprint feature to obtain the voiceprint matching score corresponding to each first voiceprint feature, which is used as the voiceprint matching score corresponding to each second audio data.

[0098] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0099] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0100] This invention provides a sound pickup and amplification method and system based on human body recognition and voiceprint matching. In scenarios with multiple speakers, it performs human body recognition on real-time acquired human images to determine the identity IDs of all personnel in the current scene. When the identity ID matches the user ID corresponding to any user with speaking authority, it controls the activation and adjusts the parameters of the sound pickup device. This ensures that the sound pickup device only collects the first audio data corresponding to the sound source location in real time when a user with speaking authority is present. This prevents invalid audio data from occupying system storage space, avoids invalid audio data from consuming system processing performance, and reduces the frequency of use of the sound pickup device, thereby extending its lifespan and fundamentally improving the overall adaptive amplification efficiency. Simultaneously, automatically adjusting the sound pickup device parameters according to the sound source location in the current scene effectively improves the sound pickup and amplification effect, enhancing the overall user experience. In addition, after the first audio data is collected, the second voiceprint feature corresponding to the user with the highest speaking authority is used to perform voiceprint recognition and matching on the first audio data. Audio data with a voiceprint matching score higher than a preset threshold is enhanced, while audio data with a voiceprint matching score lower than or equal to the preset threshold is suppressed. This achieves the gain of the speaker's audio data and the filtering of audio data other than the speaker's audio data, thereby ensuring that the speaker's audio data heard by the user is clear enough and preventing the influence of other people's audio data and ambient noise.

[0101] Furthermore, audio preservation is applied to all second audio data that has undergone enhancement or suppression processing to retain and enhance the speaker's audio data, thereby achieving further optimization of the speaker's audio data. Then, time delay correction is applied to the audio data that has undergone audio sequence preservation processing to achieve synchronized amplification and prevent unclear amplification caused by time delay.

[0102] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A sound pickup and amplification method based on human body recognition and voiceprint matching, characterized in that, include: The system acquires real-time images of the human body to be identified and the location of the sound source, and performs human body recognition on the human body image to obtain the identity IDs of all on-site personnel in the current scene. When any of the identity IDs matches the preset ID information database, the parameters of the sound pickup device in the current scene are controlled to be turned on and adjusted so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time, and perform voiceprint recognition and matching on the first audio data according to the second voiceprint features to obtain multiple second audio data and the voiceprint matching score corresponding to each second audio data. The second audio data corresponding to the voiceprint matching score that is greater than a preset threshold is enhanced, and the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold is suppressed, so as to obtain the third audio data corresponding to the first audio data, and the third audio data is amplified. The ID information database includes user IDs corresponding to all users with speaking privileges, and the second voiceprint feature is obtained by voiceprint recognition of the speaker's pre-stored voice, where the speaker is the user with the highest speaking privileges. This also includes: acquiring first action information corresponding to the speaker in real time; acquiring second action information corresponding to all on-site personnel in real time when the first action information matches a preset first action signal or the first audio data matches a preset voice signal; displaying the identity ID of the on-site personnel corresponding to the current second action information on the display screen in real time when the second action information matches a preset second action signal; and then, upon receiving a speaking right granting signal sent by the display screen, controlling and adjusting the parameters of the sound pickup device according to the speaking right granting signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time, and amplifying the fifth audio data.

2. The sound pickup and amplification method based on human body recognition and voiceprint matching as described in claim 1, characterized in that, The process involves enhancing the second audio data corresponding to the voiceprint matching score that is greater than a preset threshold, and suppressing the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold, to obtain the third audio data corresponding to the first audio data. The third audio data is then amplified. Specifically: Determine whether the voiceprint matching score of each voiceprint is greater than a preset threshold; If so, then the second audio data corresponding to the current voiceprint matching score is enhanced. If not, then the second audio data corresponding to the current voiceprint matching score is suppressed; All the second audio data that have undergone enhancement or suppression processing are subjected to audio order preservation processing to obtain the fourth audio data corresponding to the first audio data. The fourth audio data is subjected to time delay correction processing by a synchronization device to obtain the third audio data corresponding to the first audio data, and the third audio data is amplified; wherein, the synchronization device includes a time delayer, a filter and a generator.

3. The sound pickup and amplification method based on human body recognition and voiceprint matching as described in claim 1, characterized in that, The step of performing voiceprint recognition and matching on the first audio data based on the second voiceprint features to obtain multiple second audio data and the voiceprint matching score corresponding to each second audio data is as follows: Feature extraction is performed on the first audio data to obtain the audio features corresponding to the first audio data. Based on the audio features, the first audio data is segmented and clustered to obtain multiple sets of first voiceprint features and the second audio data corresponding to each first voiceprint feature. Each of the first voiceprint features and the second voiceprint features is matched to obtain the voiceprint matching score corresponding to each of the first voiceprint features, which is then used as the voiceprint matching score corresponding to each of the second audio data.

4. The sound pickup and amplification method based on human body recognition and voiceprint matching as described in claim 1, characterized in that, The process involves real-time acquisition of the human image and sound source location to be identified, and performing human recognition on the human image to obtain the identity IDs of all personnel in the current scene. Using a human body tracking algorithm from a camera, the image of the human body to be identified is captured in real time, and the location of the sound source is obtained in real time. The human image is subjected to posture recognition, and the recognized posture recognition result is matched with the pre-stored user posture information to obtain the identity IDs corresponding to all the on-site personnel in the current scene.

5. The sound pickup and amplification method based on human body recognition and voiceprint matching as described in claim 1, characterized in that, Also includes: When all the voiceprint matching scores are less than or equal to the preset threshold, a prompt signal is sent to the display screen to remind the speaker to speak.

6. A sound pickup and amplification system based on human body recognition and voiceprint matching, characterized in that, include: The data acquisition module is used to acquire the human image and sound source location to be identified in real time, and to perform human recognition on the human image to obtain the identity IDs of all on-site personnel in the current scene. The identification and matching module is used to control the activation and adjustment of the parameters of the sound pickup device in the current scene when any of the identity IDs matches the preset ID information database. This enables the sound pickup device to collect first audio data corresponding to the sound source location in real time, and to perform voiceprint recognition and matching on the first audio data based on the second voiceprint features, thereby obtaining multiple second audio data and a voiceprint matching score corresponding to each second audio data. The ID information database includes user IDs corresponding to all users with speaking permissions, and the second voiceprint features are obtained by voiceprint recognition of the speaker's pre-stored voice. The speaker is the user with the highest speaking permissions. The first amplification module is used to enhance the second audio data corresponding to the voiceprint matching score that is greater than a preset threshold, and to suppress the second audio data corresponding to the voiceprint matching score that is less than or equal to the preset threshold, so as to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data. The system also includes a second amplification module, used to acquire first action information corresponding to the speaker in real time; when the first action information matches a preset first action signal or the first audio data matches a preset voice signal, it acquires second action information corresponding to all the on-site personnel in real time; when the second action information matches a preset second action signal, it displays the identity ID of the on-site personnel corresponding to the current second action information on the display screen in real time; and when it receives a speaking right granting signal sent by the display screen, it controls and adjusts the parameters of the sound pickup device according to the speaking right granting signal, so that the sound pickup device can collect the fifth audio data of the on-site personnel corresponding to the current second action information in real time, and amplify the fifth audio data.

7. The sound pickup and amplification system based on human body recognition and voiceprint matching as described in claim 6, characterized in that, The first amplification module specifically includes: An audio adjustment unit is used to determine whether each of the voiceprint matching scores is greater than a preset threshold; if so, it enhances the second audio data corresponding to the current voiceprint matching score; if not, it suppresses the second audio data corresponding to the current voiceprint matching score. An audio order preservation unit is used to perform audio order preservation processing on all the second audio data that has undergone enhancement processing or suppression processing, so as to obtain the fourth audio data corresponding to the first audio data. A delay correction unit is used to perform delay correction processing on the fourth audio data through a synchronization device to obtain the third audio data corresponding to the first audio data, and to amplify the third audio data; wherein, the synchronization device includes a delay unit, a filter, and a generator.

8. The sound pickup and amplification system based on human body recognition and voiceprint matching as described in claim 6, characterized in that, Also includes: The reminder module is used to send a reminder signal to the display screen when all the voiceprint matching scores are less than or equal to the preset threshold, so as to remind the speaker to speak.

9. A sound pickup and amplification system based on human body recognition and voiceprint matching as described in claim 6, characterized in that, The identification and matching module specifically includes: The device control unit is used to control the activation and adjustment of the parameters of the sound pickup device in the current scene when any of the identity IDs matches the ID information database, so that the sound pickup device can collect the first audio data corresponding to the sound source location in real time; The feature extraction unit is used to extract features from the first audio data, obtain audio features corresponding to the first audio data, and perform segmentation and clustering processing on the first audio data according to the audio features to obtain multiple sets of first voiceprint features and second audio data corresponding to each first voiceprint feature. The voiceprint matching unit is used to match each of the first voiceprint features with the second voiceprint features to obtain the voiceprint matching score corresponding to each of the first voiceprint features, which is used as the voiceprint matching score corresponding to each of the second audio data.