Sound source localization method and system based on fusion confidence
Through the sound source positioning method of microphone array and voiceprint feature matching combined with face recognition, the camera scanning range is dynamically adjusted, and the three-dimensional positioning of sound source positioning and error problems in the scenes of multi-person sound production in the prior art are solved, achieving higher accuracy and reliable sound source positioning.
Patent Information
- Application Number
- CN202510835741.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing sound source positioning technology has spatial dimension defects, which cannot achieve three-dimensional positioning. In the scenario of low signal-to-noise ratio and multiple people simultaneous sounding, the positioning error is large, and the lack of a multi-modal data verification mechanism is possible, resulting in inaccurate positioning.
The sound source positioning method based on fusion confidence is adopted, and the sound source azimuth angle is calculated through the microphone array, combined with voiceprint feature matching and face recognition, the camera scanning range is dynamically adjusted, multi-modal feature correction and cross-modal verification are achieved, and positioning accuracy is improved.
While reducing the complexity of hardware, it improves the accuracy and availability of sound source positioning, solves the accuracy problems in three-dimensional positioning and noisy environments, and improves the user experience.
Smart Images

Figure CN120352834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sound source localization, and provides a sound source localization method and system based on fused confidence. Background Art
[0002] Sound Source Localization (SSL) is a technology that determines the spatial position of a sound source by analyzing the characteristics of sound wave propagation, collecting signals using a multi-microphone array, and calculating information such as time difference and phase difference. Its core objective is to restore the direction (azimuth angle, elevation angle) or distance parameters of the sound source through mathematical modeling and signal processing. Based on the time difference, sound level difference, or phase difference of the sound wave arriving at different microphones, equations are constructed in combination with geometric relationships to solve the position of the sound source. For example, a triangular microphone array can calculate the time delay through the cross-correlation algorithm and reverse the direction of the sound source.
[0003] The following problems exist in the prior art: spatial dimension defect, the dual-microphone system is limited by geometric constraints and can only perform planar localization, unable to resolve the elevation angle and front-back azimuth of the sound source; single-modal defect: the positioning error of the pure acoustic method increases when the SNR is less than 10 dB, for example, it can reach 6 cm to 8 cm; scene adaptability defect, lacking a multi-modal data verification mechanism, unable to solve the positioning conflicts in scenarios such as sound source movement and multiple people speaking simultaneously.
[0004] Therefore, it is necessary to provide a new sound source localization method and system based on fused confidence to solve the above problems. Summary of the Invention
[0005] The present invention provides a sound source localization method and system based on fused confidence to solve the technical problems in the prior art, such as large positioning errors caused by single-mode data and the inability to accurately locate in scenarios such as sound source movement and multiple people speaking simultaneously due to the lack of a multi-modal data verification mechanism. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0006] A method for sound source localization based on fusion confidence is proposed in the first aspect of the present invention, including: when a preset wake-up word is detected, starting the sound source localization process, specifically including calculating the horizontal azimuth angle according to the detected current voice signal and determining the sound source localization area; extracting multi-dimensional voice signals from the current voice signal to generate a specified-dimensional voiceprint feature vector, calculating the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; calculating the user preference fusion coefficient corresponding to the current voice signal, and further calculating the fusion confidence of the current voice signal according to the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; dynamically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so as to capture the image frame of the current user at a fixed interval during the scanning process, perform face detection, and obtain the face image to be processed; extracting facial features from the face image to be processed for visual identity collaborative confirmation, and when the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, stopping the rotation of the pan-tilt camera and locking the current azimuth as the determined sound source position.
[0007] A sound source localization system based on fusion confidence is proposed in the second aspect of the present invention. It executes the sound source localization method described in the first aspect of the present invention. The sound source localization system includes: a pan-tilt camera, a first microphone and a second microphone located on both sides of the pan-tilt camera; a sound source localization module, configured to start the sound source localization process when a preset wake-up word is detected, specifically including calculating the horizontal azimuth angle according to the detected current voice signal and determining the sound source localization area; a voiceprint recognition module, configured to extract multi-dimensional voice signals from the current voice signal to generate a specified-dimensional voiceprint feature vector, calculate the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; a control processing unit, configured to calculate the user preference fusion coefficient corresponding to the current voice signal, and further calculate the fusion confidence of the current voice signal according to the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; a pan-tilt mechanical structure unit, dynamically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so that the visual processing module captures the image frame of the current user at a fixed interval during the scanning process, performs face detection, and obtains the face image to be processed; a determination module, configured to extract facial features from the face image to be processed for visual identity collaborative confirmation, and when the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, stopping the rotation of the pan-tilt camera and locking the current azimuth as the determined sound source position.
[0008] A third aspect of the present invention provides an electronic device, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for sound source localization based on fusion confidence described in the first aspect of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for sound source localization based on fusion confidence described in the first aspect of the present invention is implemented.
[0010] The embodiments of the present invention include the following advantages: Compared with the prior art, the present invention constructs a three-level collaborative architecture of acoustic localization, face feature recognition and visual verification, provides a positioning space reference through preliminary sound source azimuth localization, narrows down the user range by combining voiceprint feature matching, and introduces a dynamic scanning compensation mechanism, specifically automatically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence, which can effectively improve the positioning accuracy while reducing the hardware complexity.
[0011] In addition, a cross-modal cross-verification channel between the voiceprint ID and the face ID is established, compensating for the spatial perception limitations of the dual microphones through learning from historical azimuth data, forming a closed-loop system of mutual correction of multi-modal features to improve the user experience. The introduction of voiceprint recognition and face recognition technologies solves the problems of the inability of the dual microphone array to achieve three-dimensional positioning, low positioning accuracy, and low accuracy in noisy environments, and improves the usability and user experience of the sound source localization function. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of the steps of an example of the method for sound source localization based on fusion confidence of the present invention; Figure 2 is a schematic framework diagram of a specific application example of the method for sound source localization based on fusion confidence of the present invention; Figure 3 is an example diagram of the sound source localization area determined in the method for sound source localization based on fusion confidence of the present invention; Figure 4 is a structural block diagram of the sound source localization system based on fusion confidence of the present invention; Figure 5 is a schematic structural diagram of an embodiment of an electronic device according to the present invention; Figure 6 is a schematic structural diagram of an embodiment of a computer-readable medium according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are for the purpose of enabling those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.
[0014] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0015] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices. Additionally, in the accompanying drawings, the left-right direction on the drawing page is the horizontal direction, and the up-down direction on the drawing page is the vertical direction (or the direction perpendicular to the horizontal direction).
[0016] It should be understood that although the ordinal adjectives such as first, second, third, etc. may be used herein to describe various devices, elements, components or parts, this should not be limiting. These ordinal adjectives are used to distinguish one from another. For example, the first device may also be called the second device without departing from the essential technical solution of the present invention.
[0017] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0018] In view of the above problems, the present invention proposes a sound source localization method based on fusion confidence. In this method, when a preset wake-up word is detected, the sound source localization process is started, which specifically includes calculating the horizontal azimuth angle according to the detected current voice signal to determine the sound source localization area; extracting multi-dimensional voice signals from the current voice signal to generate a specified-dimension voiceprint feature vector, calculating the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; calculating the user preference fusion coefficient corresponding to the current voice signal, and further calculating the fusion confidence of the current voice signal according to the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; dynamically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so as to capture the image frame of the current user at a fixed interval during the scanning process, perform face detection, and obtain the face image to be processed; extracting facial features from the face image to be processed for visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, stop the rotation of the pan-tilt camera and lock the current azimuth as the determined sound source position. By constructing a three-level collaborative architecture of acoustic localization, face feature recognition, and visual verification, a positioning space benchmark is provided through preliminary sound source azimuth localization, the user range is narrowed by combining voiceprint feature matching, and a dynamic scanning compensation mechanism is introduced. Specifically, the scanning range of the pan-tilt camera is automatically adjusted according to the calculated fusion confidence, which can effectively improve the positioning accuracy while reducing the hardware complexity.
[0019] Embodiment 1 The following will refer to Figure 1 、 Figure 2 、 Figure 3 to describe the content of the present invention in detail.
[0020] Figure 1 is a step flowchart of an example of the sound source localization method based on fusion confidence of the present invention. Figure 2 is a framework schematic diagram of a specific application example of the sound source localization method of the present invention. Figure 3 is an example diagram of the sound source localization area determined in the sound source localization method based on fusion confidence of the present invention.
[0021] Refer to Figure 1 and Figure 2 In step S101, when a preset wake-up word is detected, the sound source localization process is started, which specifically includes calculating the horizontal azimuth angle according to the detected current voice signal to determine the sound source localization area.
[0022] In Figure 2In the application example, it includes a pan-tilt camera, a first microphone and a second microphone located on both sides of the pan-tilt camera. The first microphone and the second microphone are symmetrically arranged relative to the pan-tilt camera to form a dual microphone array.
[0023] For a voice device in a smart home scenario, when the user utters a preset wake-up word (such as "Hello **"), through the collaborative work of the sound source localization module, the voiceprint recognition module, and the visual processing module (integrated with face recognition technology), the sound source position of the current voice signal (i.e., the position of the current user) is identified, and the pan-tilt camera is automatically controlled to accurately turn towards the user, enabling the user to communicate face-to-face with the voice device naturally without moving or searching for the voice device, enhancing the convenience and experience of the interaction. Among them, the voice device includes a speaker, a TV box, and a pan-tilt camera.
[0024] Specifically, when a preset wake-up word (such as "Hello **") is detected, sound source localization processing is started, which specifically includes calculating the horizontal azimuth angle based on the detected voice signal and determining the sound source localization area.
[0025] The following expression is used to calculate the horizontal azimuth angle of the current voice signal: θ = arcsin(c * Δt / d) (1) Where θ represents the angle formed by the sound source direction of the current voice signal and the central axis perpendicular to the horizontal direction of the pan-tilt camera, also known as the horizontal azimuth angle, θ∈[-90°,90°]; c represents the sound propagation speed, with the unit of meters per second; d represents the distance between the first microphone and the second microphone in the horizontal direction, with the unit of meters, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal.
[0026] In this example, the value of d is preferably 8 meters. Preferably, the horizontal azimuth angle θ∈[-75°,75°].
[0027] It should be noted that through historical detection data, it is determined that when the horizontal azimuth angle θ∈[-75°,75°], the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal can be measured more accurately, and the calculation error is small. It can also provide a more reliable data basis for subsequent voiceprint feature extraction and fusion confidence evaluation.
[0028] According to the calculated horizontal azimuth angle, the sound source localization area is determined, with the central axis of the pan-tilt camera (such as Figure 3The Z-axis line represented by a dashed line in [description] (the center of the pan-tilt camera is point O) towards the area range that forms a horizontal azimuth angle approaching the first microphone, and the area range that forms a horizontal azimuth angle (such as "θ") approaching the second microphone along the central axis of the pan-tilt camera, to be used as the sound source localization area. For details, see Figure 3 the "OHL area" shown in
[0029] When the signal-to-noise ratio of the current voice signal is greater than or equal to the set threshold, the sound source localization score of the current voice signal is greater than or equal to 0.8.
[0030] Optionally, the set threshold is 20dB - 25dB. In this example, the set threshold is 20dB.
[0031] It should be noted that the above is described as an optional example and should not be construed as a limitation to the present invention.
[0032] Next, in step S102, multi-dimensional voice signal extraction is performed on the current voice signal to generate a specified-dimensional voiceprint feature vector, and the existing user voiceprint that matches the current voice signal is calculated to confirm the current user corresponding to the current voice signal.
[0033] Using Mel Frequency Cepstral Coefficients, extract the specific-dimensional voice features of the current voice signal, and generate a multi-dimensional voiceprint feature vector through a pre-trained model. The dimension of the multi-dimensional voiceprint feature vector is greater than the dimension of the specific-dimensional voice features.
[0034] In a specific embodiment, Mel Frequency Cepstral Coefficients (MFCC) are used to extract voice features of, for example, 39 dimensions of the current voice signal, that is, specific-dimensional voice features. The voice features of 39 dimensions include static spectrum information and dynamic change information, specifically including twelve cepstral coefficients, energy, first-order difference, and second-order difference corresponding to the current voice signal.
[0035] Based on a CNN network, it is trained using a training data set to obtain a pre-trained model for generating a multi-dimensional voiceprint feature vector corresponding to the current voice signal. The pre-trained model includes an input layer, five convolutional layers, three fully connected layers, and an output layer. The training data set includes voice data (including voice signals) labeled with user voiceprint identifiers (such as user voiceprint IDs). In the multi-classification task of model training, the voice data labeled with the target user identifier (such as user voiceprint ID) is used as the positive sample, and the voice data of the remaining users are used as negative samples.
[0036] Specifically, the current voice signal and the specific dimension voice features corresponding to the current voice signal are input into a pre-trained model to obtain a multi-dimensional voiceprint feature vector corresponding to the current voice signal. The multi-dimensional voiceprint feature vector is, for example, a 128-dimensional voiceprint feature vector.
[0037] Furthermore, similarity matching calculation is performed between the calculated multi-dimensional voiceprint feature vector and the existing user voiceprints in the feature database. Specifically, the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal is calculated to be used to determine the existing user voiceprint that matches.
[0038] For the similarity matching calculation, when the calculated matching degree is greater than the matching degree threshold (for example, greater than or equal to 0.85), the existing user voiceprint that matches the current voice signal is determined. When the calculated matching degree is less than the matching degree threshold (for example, less than 0.85), the next existing user voiceprint is continuously queried from the feature database, and the above similarity matching calculation is performed until the existing user voiceprint that matches the current voice signal is determined.
[0039] For the feature database, each user is guided to read the wake-up word to form a voice signal. The voice device collects an audio segment with a signal-to-noise ratio greater than or equal to 25 dB, for example, and extracts the Mel-frequency cepstral features of the formed voice signal. The pre-trained model is used to generate a voiceprint feature vector corresponding to the formed voice signal. Each voiceprint feature vector is bound to the corresponding user account to generate binding information, and the binding information is stored to establish the feature database. Or, by recording the user voiceprints of each user, the user voiceprint identifiers of each user are extracted, and the user voiceprint identifiers and user accounts are formed into binding information, and the binding information is stored to establish the feature database.
[0040] For example, the feature database includes user information, user account information (such as user account, user ID, etc.), one or more voiceprint feature vectors corresponding to each user account, etc.
[0041] It should be noted that the above is described as an optional example and should not be construed as a limitation to the present invention.
[0042] Next, in step S103, the user preference fusion coefficient corresponding to the current voice signal is calculated, and based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal, the fusion confidence of the current voice signal is further calculated.
[0043] Generate an azimuth correction coefficient based on the user's azimuth preference data in the historical interaction data between the user and the home device (for example, the mean value of the common angle distribution of the horizontal azimuth angle in the historical time periods such as one month, two months, three months, six months, and twelve months calculated backward from the current time), and the proportion of the actual user default events counted in all interaction samples, that is, the default rate. Specifically, use the following expression to calculate the azimuth correction coefficient: α = q×(1 - δ / δ max )(2) Among them, α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles in a specific time period, and α∈[0.05, 0.2]; q represents the optimal coefficient, which is determined by performing regression analysis and fitting on the time delay data and / or intensity data at different angles; δ max represents the historical maximum value in the proportion of the actual user default events counted in all interaction samples, that is, the historical highest default rate; δ represents the default rate of the user when the position where the user wakes up the voice device is not within the specified position.
[0044] It should be noted that in the present invention, the value of q in the above expression (2) is preferably 0.15. The optimal coefficient q is obtained by calculating through data regression analysis. Specifically, 200 groups of time delay data and / or intensity data at different angles in the home scene are collected, and the optimal coefficient obtained by performing regression analysis and fitting is used to minimize the positioning error. In addition, the actual user default event refers to the situation where the position where the user wakes up the voice device is not within the specified position (for example, not within the angular range formed by the specified horizontal angle. The specified horizontal angle is in the range of -75° to 75°, and optionally, in the range of -77° to 73°. When not within the range formed by the above specified horizontal angle, it is an actual user default event. The above is only an optional example for illustration and should not be construed as a limitation to the present invention.
[0045] Next, further calculate the fusion confidence of the current voice signal based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal, the signal-to-noise ratio determined by the current voice signal, and the sound source localization score generated by the time difference, including: Calculate the fusion confidence of the current voice signal using the following expression: C fusion = k*S loc + w*S voice + α*Δθ (3) Among them, C fusion represents the fusion confidence of the current voice signal; S locThe sound source localization score generated based on the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the first microphone and the second microphone; S voice represents the calculated matching degree between the current voice signal and the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth angle and the average value of the historical horizontal azimuth angles within a specific time period, Δθ = |θ - μ|, where θ represents the current horizontal azimuth angle and μ represents the average value of the historical horizontal azimuth angle distribution within a specific time period, that is, the average value of the historical horizontal azimuth angles; k represents the first parameter, specifically the first parameter corresponding to the sound source localization score generated based on the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the dual microphones; w represents the second parameter, specifically the second parameter corresponding to the calculated matching degree between the current voice signal and the existing user voiceprint; α represents the azimuth correction coefficient, specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the average value of the historical horizontal azimuth angles within a specific time period.
[0046] Optionally, the value range of k is 0.55 to 0.75, the value range of w is 0.25 to 0.45, and the following expression is satisfied: k + w = 1.
[0047] Preferably, the value of k is 0.6, and the value range of w is 0.4.
[0048] It should be noted that for k with a value of 0.6 and w with a value range of 0.4, the optimal coefficients are obtained by collecting 200 groups of time delay data and / or intensity data at different angles in the home scene and using regression analysis fitting to minimize the positioning error. The preferred values of k and w are the values corresponding to the optimal coefficients and are determined based on the optimal coefficients and the minimum positioning error. The above is described as an optional example and should not be construed as a limitation to the present invention.
[0049] Next, in step S104, according to the calculated fusion confidence of the current voice signal, dynamically adjust the scanning range of the pan-tilt camera to capture the image frame of the current user at a fixed interval during the scanning process, perform face detection, and obtain the face image to be processed.
[0050] Specifically, dynamically adjust the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal.
[0051] When the calculated fusion confidence of the current voice signal is greater than or equal to 0.9, increase the scanning angle of the pan-tilt camera by 5° or decrease it by 5°.
[0052] When the calculated fusion confidence of the current voice signal is greater than or equal to 0.7 and less than 0.9, increase the scanning angle of the pan-tilt camera by 10° or decrease it by 10°.
[0053] When the fusion confidence of the calculated current voice signal is less than 0.7, increase the scanning angle of the pan-tilt camera by 15° or decrease it by 15°.
[0054] Further, during the scanning process of the pan-tilt camera, capture image frames of the current user at fixed intervals, perform face detection, and obtain face images to be processed.
[0055] The fixed interval is, for example, 0.1s to 0.5s.
[0056] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.
[0057] Next, in step S105, perform facial feature extraction on the face image to be processed for visual identity collaborative verification. When the visual identity collaborative verification result of the face image to be processed meets the identity consistency condition, stop the rotation of the pan-tilt camera and lock the current orientation, that is, determine the sound source position.
[0058] Specifically, perform facial feature extraction on the face image to be processed for visual identity collaborative verification. When the visual identity collaborative verification result of the face image to be processed meets the identity consistency condition, stop the rotation of the pan-tilt camera and lock the current orientation as the determined sound source position.
[0059] Specifically, perform facial feature extraction on the face image to be processed. Use the FaceNet model to extract, for example, a 512-dimensional feature vector. The 512-dimensional feature vector includes geometric features such as eye distance and nose tip coordinates, texture features (local binary pattern LBP), and global features (facial contour encoding).
[0060] For visual identity collaborative verification, perform cross-modal verification. Specifically, compare whether the user identifier associated with the face feature (such as UserID) is consistent with the user identifier corresponding to the voiceprint feature vector of the current voice information (such as UserID). Specifically, according to a consistency threshold (such as 0.92), or use the Euclidean distance to determine consistency (such as by determining that when the calculated Euclidean distance is less than or equal to 1.2) to determine whether the user identifier associated with the face feature (such as UserID) is consistent with the user identifier corresponding to the voiceprint feature vector of the current voice information.
[0061] In a specific embodiment, the Euclidean distance between the user identifier corresponding to the human body image to be processed and the user identifier corresponding to the voiceprint feature vector of the current voice information is specifically calculated. When the calculated Euclidean distance is less than or equal to 1.2, it indicates that the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition. When the calculated Euclidean distance is greater than 1.2, it indicates that the visual identity collaborative confirmation result of the face image to be processed does not meet the identity consistency condition.
[0062] In another specific embodiment, the similarity between the user identifier associated with the face feature (such as UserID) and the user identifier corresponding to the voiceprint feature vector of the current voice information (such as UserID) is calculated. When the calculated similarity is greater than the consistency threshold (such as 0.92), it indicates that the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition. When the calculated similarity is less than or equal to the consistency threshold (such as 0.92), it indicates that the visual identity collaborative confirmation result of the face image to be processed does not meet the identity consistency condition.
[0063] When the identity consistency condition is met, the control processing unit controls the pan-tilt mechanical structure unit to immediately stop the rotation of the pan-tilt camera in the horizontal direction and lock the current orientation. Specifically, the center of the camera is facing the current user, and the locked current orientation is used as the determined sound source position.
[0064] It should be noted that the consistency threshold is determined by optimizing the test set of the cross-modal verification model. Specifically, on the premise of ensuring a 98% true acceptance rate (TAR), the false acceptance rate (FAR) is reduced to less than or equal to 0.1%. Under the above limited conditions, the determined consistency threshold is preferably 0.92. The above is only for illustrative purposes and should not be construed as a limitation of the present invention.
[0065] Compared with the prior art, the present invention constructs a three-level collaborative architecture of acoustic positioning, face feature recognition, and visual verification. The acoustic source azimuth is initially positioned to provide a positioning space reference, the user range is narrowed by combining voiceprint feature matching, and a dynamic scanning compensation mechanism is introduced. Specifically, the scanning range of the pan-tilt camera is automatically adjusted according to the calculated fusion confidence, effectively improving the positioning accuracy while reducing the hardware complexity.
[0066] In addition, a cross-modal cross-verification channel between the voiceprint ID and the face ID is established. The spatial perception limitation of the dual microphone is compensated by learning historical azimuth data, forming a closed-loop system for mutual correction of multi-modal features, improving the user experience. By introducing voiceprint recognition and face recognition technologies, the problems of the dual microphone array being unable to achieve three-dimensional positioning, low positioning accuracy, and low accuracy in noisy environments are solved, improving the usability of the sound source positioning function and the user experience.
[0067] Example 2 The following is a system embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the system embodiment of the present invention, please refer to the method embodiment of the present invention.
[0068] Figure 4 is a schematic structural diagram of an example of a sound source localization system based on fusion confidence according to the present invention. The following will refer to Figure 4 , and the sound source localization system 400 based on fusion confidence will be described. The sound source localization system 400 executes the sound source localization method described in Embodiment 1 of the present invention.
[0069] Specifically, the sound source localization system 400 includes a pan-tilt camera 410, a first microphone 420 and a second microphone 430 located on both sides of the pan-tilt camera, a sound source localization module 440, a voiceprint recognition module 450, a control processing unit 460, a pan-tilt mechanical structure unit 470, and a determination module 480.
[0070] Further, the sound source localization module 440 is used to start sound source localization processing when a preset wake-up word is detected, specifically including calculating a horizontal azimuth angle according to the detected current voice signal and determining a sound source localization area. The voiceprint recognition module 450 is used to extract multi-dimensional voice signal features from the current voice signal to generate a specified dimension voiceprint feature vector, and calculate an existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal. The control processing unit 460 is used to calculate a user preference fusion coefficient corresponding to the current voice signal, and further calculate a fusion confidence of the current voice signal according to a matching degree obtained by calculating an existing user voiceprint that matches the current voice signal. The pan-tilt mechanical structure unit 470 dynamically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so that the visual processing module captures an image frame of the current user at a fixed interval during the scanning process, performs face detection, and obtains a face image to be processed. The determination module 480 is used to extract facial features from the face image to be processed for visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the rotation of the pan-tilt camera is stopped, and the current orientation is locked as the determined sound source position.
[0071] Optionally, the sound source localization system is used in cooperation with an audio device. Specifically, the sound source localization module is used to monitor a preset wake-up word for the current user to wake up the audio device. The audio device includes an audio and a TV box.
[0072] It should be noted that in other embodiments, the audio device is an integrated product, specifically including an audio, a TV box, and a camera (such as a pan-tilt camera). The above is only described as an optional example and should not be construed as a limitation to the present invention.
[0073] According to an optional embodiment, based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal, the fusion confidence level of the current voice signal is further calculated.
[0074] The fusion confidence level of the current voice signal is calculated using the following expression: C fusion =k*S loc +w*S voice +α*Δθ where C fusion represents the fusion confidence level of the current voice signal; S loc is the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the first microphone and the second microphone; S voice represents the matching degree of the calculated current voice signal with the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles within a specific time period, Δθ = ∣θ - μ∣, θ represents the current horizontal azimuth angle, and μ represents the mean value of the historical horizontal azimuth angle distribution within a specific time period, that is, the historical horizontal azimuth angle mean value; k represents a first parameter, specifically the first parameter corresponding to the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the two microphones; w represents a second parameter, specifically the second parameter corresponding to the matching degree of the calculated current voice signal with the existing user voiceprint; α represents an azimuth correction coefficient, specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles within a specific time period.
[0075] According to an optional embodiment, based on the user azimuth preference data in the historical interaction data between the user and the home device and the proportion of the statistically actual user default events in all interaction samples, that is, the default rate, an azimuth correction coefficient is generated. Specifically, the following expression is used to calculate the azimuth correction coefficient: α=q×(1-δ / δ max ) where α represents the azimuth correction coefficient, specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles within a specific time period, α ∈ [0.05, 0.2]; q represents the optimal coefficient, which is determined by performing regression analysis and fitting based on the time delay data and / or intensity data at different angles; δ maxIt represents the historical maximum value in the proportion of the actual user default events counted in all interaction samples, that is, the historical highest default rate; δ represents the user default rate when the position where the current user wakes up the voice device is not within the specified position.
[0076] The value range of k is 0.55 to 0.75, and the value range of w is 0.25 to 0.45, and it satisfies the following expression: k + w = 1.
[0077] The following expression is used to calculate the horizontal azimuth angle of the current voice signal: θ = arcsin(c * Δt / d) Where, θ represents the included angle formed by the sound source direction of the current voice signal and the central axis perpendicular to the horizontal direction of the pan-tilt camera, θ ∈ [-90°, 90°]; c represents the sound propagation speed, with the unit of meter per second; d represents the horizontal distance between the first microphone and the second microphone, with the unit of meter, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal.
[0078] According to the calculated horizontal azimuth angle, determine the sound source positioning area, that is, the area range formed by the central axis of the pan-tilt camera towards the first microphone to form a horizontal azimuth angle, and the area range formed by the central axis of the pan-tilt camera towards the second microphone to form a horizontal azimuth angle, as the determined sound source position.
[0079] According to the optional implementation manner, when the signal-to-noise ratio of the current voice signal is greater than or equal to the set threshold, the sound source positioning score of the current voice signal is greater than or equal to 0.8, and the set threshold is 20dB to 25dB.
[0080] According to the optional implementation manner, according to the calculated fusion confidence of the current voice signal, dynamically adjust the scanning range of the pan-tilt camera.
[0081] When the calculated fusion confidence of the current voice signal is greater than or equal to 0.9, increase or decrease the scanning angle of the pan-tilt camera by 5°.
[0082] When the calculated fusion confidence of the current voice signal is greater than or equal to 0.7 and less than 0.9, increase or decrease the scanning angle of the pan-tilt camera by 10°.
[0083] When the calculated fusion confidence of the current voice signal is less than 0.7, increase or decrease the scanning angle of the pan-tilt camera by 15°.
[0084] Using Mel-frequency cepstral coefficients, specific-dimensional voice features of the current voice signal are extracted, and multi-dimensional voiceprint feature vectors are generated through a pre-trained model. The dimension of the multi-dimensional voiceprint feature vectors is greater than that of the specific-dimensional voice features.
[0085] Preferably, It should be noted that since Figure 4 the sound source localization method performed by the sound source localization system based on the fusion confidence is substantially the same as Figure 1 the sound source localization method in the example of , the description of the same part is omitted.
[0086] Compared with the prior art, the present invention constructs a three-level collaborative architecture of acoustic localization, face feature recognition and visual verification. By initially localizing the sound source azimuth to provide a localization space reference, combining voiceprint feature matching to narrow down the user range, and introducing a dynamic scanning compensation mechanism, specifically automatically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence, which can effectively improve the localization accuracy while reducing the hardware complexity.
[0087] In addition, a cross-modal cross-verification channel between the voiceprint ID and the face ID is established. By learning from historical azimuth data to compensate for the spatial perception limitations of the dual microphones, a closed-loop system for mutual correction of multi-modal features is formed to improve the user experience. By introducing voiceprint recognition and face recognition technologies, the problems of inability to achieve three-dimensional localization, low localization accuracy, and low accuracy in noisy environments of the dual microphone array are solved, and the usability of the sound source localization function and the user experience are improved.
[0088] Figure 5 It is a schematic structural diagram of an embodiment of an electronic device according to the present invention.
[0089] As Figure 5 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work collaboratively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0090] The memory stores computer-executable programs, usually machine-readable code. The computer-readable program can be executed by the processor so that the electronic device can execute the method of the present invention or at least part of the steps in the method.
[0091] The memory includes volatile memory, such as a random access storage unit (RAM) and / or a cache storage unit, and can also be non-volatile memory, such as a read-only storage unit (ROM).
[0092] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface may represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the multiple bus structures.
[0093] It should be understood that Figure 5 the electronic device shown is merely an example of the present invention, and the electronic device of the present invention may further include elements or components not shown in the above example. For example, some electronic devices further include a display unit such as a display screen, and some electronic devices further include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute the computer-readable program in the memory to implement at least part of the steps of the method of the present invention, it can be considered as the electronic device covered by the present invention.
[0094] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, as Figure 6 shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.
[0095] The software product may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0096] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, and the readable medium may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0097] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0098] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the computer-readable medium implements the data interaction method of the present disclosure.
[0099] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are only different from the present embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0100] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which may be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.
[0101] It should be noted that the above detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application pertains.
[0102] In the foregoing detailed description, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.
[0103] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A sound source localization method based on fusion confidence, characterized in that, Including: When a preset wake-up word is detected, start sound source localization processing, specifically including calculating the horizontal azimuth angle based on the detected current voice signal and determining the sound source localization area; Extract multi-dimensional voice signals from the current voice signal to generate a specified-dimensional voiceprint feature vector, and calculate the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; Calculate the user preference fusion coefficient corresponding to the current voice signal, and further calculate the fusion confidence of the current voice signal based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; Dynamically adjust the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, capture image frames of the current user at fixed intervals during the scanning process, perform face detection, and obtain the face image to be processed; Extract facial features from the face image to be processed for visual identity collaborative verification. When the visual identity collaborative verification result of the face image to be processed meets the identity consistency condition, stop the rotation of the pan-tilt camera and lock the current orientation as the determined sound source position.
2. The sound source localization method according to claim 1, wherein The further calculation of the fusion confidence of the current voice signal based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal includes: Calculate the fusion confidence of the current voice signal using the following expression: C fusion = k * S loc + w * S voice + α * Δθ; Among them, C fusion represents the fusion confidence of the current voice signal; S loc is the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the first microphone and the second microphone; S voice represents the matching degree between the calculated current voice signal and the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth angle and the average value of the historical horizontal azimuth angles within a specific time period, Δθ = ∣θ - μ∣, where θ represents the current horizontal azimuth angle and μ represents the average value of the historical horizontal azimuth angle distribution within a specific time period, that is, the average value of the historical horizontal azimuth angles; k represents the first parameter, specifically the first parameter corresponding to the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the two microphones; w represents the second parameter, specifically the second parameter corresponding to the matching degree between the calculated current voice signal and the existing user voiceprint; α represents the azimuth correction coefficient, specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the average value of the historical horizontal azimuth angles within a specific time period.
3. The sound source localization method according to claim 2, wherein, Further including: Generate an azimuth correction coefficient based on the user azimuth preference data in the historical interaction data between the user and the home device and the proportion of the statistically actual user default events in all interaction samples, that is, the default rate. Specifically, use the following expression to calculate the azimuth correction coefficient: α = q×(1 - δ / δ max ); Among them, α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the average value of the historical horizontal azimuth angles within a specific time period, and α ∈ [0.05, 0.2]; q represents the optimal coefficient, which is determined by performing regression analysis and fitting based on the time delay data and / or intensity data at different angles; δ max represents the historical maximum value in the proportion of the actual default events of the statistically users in all interaction samples, that is, the historical highest default rate; δ represents the default rate of the current user whose position of waking up the voice device is not within the specified position.
4. The sound source localization method according to claim 2, characterized in that, Further including: The value range of k is 0.55 to 0.75, the value range of w is 0.25 to 0.45, and the following expression is satisfied: k + w = 1.
5. The sound source localization method according to claim 2, characterized in that, Further including: Use the following expression to calculate the horizontal azimuth angle of the current voice signal: θ = arcsin(c * Δt / d); Where, θ represents the angle formed by the sound source direction of the current voice signal and the central axis perpendicular to the horizontal direction of the pan-tilt camera, θ ∈ [-90°, 90°]; c represents the sound propagation speed, in meters per second; d represents the distance between the first microphone and the second microphone in the horizontal direction, in meters, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal; Determine the sound source localization area according to the calculated horizontal azimuth angle, and use the area range formed by the central axis of the pan-tilt camera approaching the first microphone to form a horizontal azimuth angle, and the area range formed by the central axis of the pan-tilt camera approaching the second microphone to form a horizontal azimuth angle as the determined sound source position.
6. The sound source localization method according to claim 5, characterized in that, Further including: When the signal-to-noise ratio of the current voice signal is greater than or equal to the set threshold, the sound source localization score of the current voice signal is greater than or equal to 0.8, and the set threshold is 20 dB to 25 dB.
7. The sound source localization method according to claim 1, wherein Dynamically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal includes: When the calculated fusion confidence of the current voice signal is greater than or equal to 0.9, increase or decrease the scanning angle of the pan-tilt camera by 5°; When the calculated fusion confidence of the current voice signal is greater than or equal to 0.7 and less than 0.9, increase or decrease the scanning angle of the pan-tilt camera by 10°; When the calculated fusion confidence of the current voice signal is less than 0.7, increase or decrease the scanning angle of the pan-tilt camera by 15°.
8. The sound source localization method according to claim 1, wherein Further includes: Using Mel Frequency Cepstral Coefficients to extract specific-dimensional voice features of the current voice signal, and generating a multi-dimensional voiceprint feature vector through a pre-trained model, where the dimension of the multi-dimensional voiceprint feature vector is greater than the dimension of the specific-dimensional voice features.
9. A sound source localization system based on fusion confidence, characterized in that, It executes the sound source localization method described in any one of claims 1 to 8, and the sound source localization system includes: A pan-tilt camera, a first microphone and a second microphone located on both sides of the pan-tilt camera; A sound source localization module for starting sound source localization processing when a preset wake-up word is detected, specifically including calculating the horizontal azimuth angle according to the detected current voice signal and determining the sound source localization area; A voiceprint recognition module for extracting multi-dimensional voice signals from the current voice signal to generate a specified-dimensional voiceprint feature vector, and calculating the existing user voiceprint matching the current voice signal to confirm the current user corresponding to the current voice signal; A control processing unit for calculating a user preference fusion coefficient corresponding to the current voice signal, and further calculating the fusion confidence of the current voice signal according to the matching degree obtained by calculating the existing user voiceprint matching the current voice signal; A pan-tilt mechanical structure unit dynamically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so that the visual processing module captures image frames of the current user at fixed intervals during the scanning process, performs face detection, and obtains a face image to be processed; A determination module for extracting facial features from the face image to be processed for visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, stop the rotation of the pan-tilt camera and lock the current orientation as the determined sound source position.
10. The sound source localization system according to claim 9, characterized in that, Further includes: Calculated using the following expression, the fusion confidence of the current voice signal: C fusion = k * S loc + w * S voice + α * Δθ; Among them, C fusion represents the fusion confidence of the current voice signal; S loc is the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the first microphone and the second microphone; S voice represents the matching degree between the calculated current voice signal and the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles within a specific time period, Δθ = ∣θ - μ∣, where θ represents the current horizontal azimuth angle and μ represents the mean value of the historical horizontal azimuth angle distribution within a specific time period, that is, the historical horizontal azimuth angle mean value; k represents the first parameter, specifically the first parameter corresponding to the sound source localization score generated from the signal-to-noise ratio and time difference of arrival determined when collecting the current voice signal through the two microphones; w represents the second parameter, specifically the second parameter corresponding to the matching degree between the calculated current voice signal and the existing user voiceprint; α represents the azimuth correction coefficient, specifically used to correct the absolute deviation value between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angles within a specific time period.
Citation Information
Patent Citations
Intelligent robot rotation method based on sound source positioning and face detection
CN106292732A
Object recognition method and device, storage medium and terminal
CN108305615A
Sound source positioning method and audio equipment
CN113640744A
Bimodal identity authentication method and device and storage medium
CN114398611A
User identification system through sound localization based audio-visual under robot environments and method thereof
KR100822880B1
Cited By
Face detection real-time synchronous pickup data processing method and system
CN121284291A
Face detection real-time synchronous sound pickup data processing method and system
CN121284291B