A sound source localization method and system based on fusion confidence
Through the three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification, combined with voiceprint feature matching and dynamic scanning compensation mechanism, the positioning error problems of three-dimensional positioning and low signal-to-noise ratio in existing sound source localization technology are solved, and higher-precision and reliable sound source localization is achieved.
Patent Information
- Application Number
- CN202510835741.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing sound source localization technology has spatial dimension defects and cannot achieve three-dimensional positioning. In addition, the positioning error is large in low signal-to-noise ratio and noisy environments. The lack of multimodal data verification mechanism makes it impossible to resolve positioning conflicts in scenarios such as sound source movement and multiple people speaking at the same time.
A sound source localization method based on fusion confidence is adopted. Through a three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification, combined with voiceprint feature matching and dynamic scanning compensation mechanism, the microphone array is used to calculate the sound source azimuth, combined with user preferences and historical interaction data, the camera scanning range is dynamically adjusted, and multimodal feature correction is performed.
It improves the accuracy and availability of sound source positioning, solves the problem of dual-microphone array being unable to achieve three-dimensional positioning and positioning error under low signal-to-noise ratio, and enhances user experience.
Smart Images

Figure CN120352834B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sound source localization, and provides a sound source localization method and system based on fusion confidence. Background Art
[0002] Sound source localization (SSL) is a technology that determines the spatial location of a sound source by analyzing the propagation characteristics of sound waves, using a multi-microphone array to collect signals and calculate information such as time and phase differences. Its core goal is to recover the direction (azimuth, elevation) or distance parameters of the sound source through mathematical modeling and signal processing. Based on the time difference, sound level difference, or phase difference between the sound waves reaching different microphones, geometric relationships are combined to construct equations to solve for the sound source location. For example, a triangular microphone array can use a cross-correlation algorithm to calculate time delay and infer the direction of the sound source.
[0003] The existing technology has the following problems: spatial dimensionality defects. The dual-microphone system is limited by geometric constraints and can only perform planar positioning, but cannot resolve the elevation angle and front-back direction of the sound source; single modality defects: the positioning error of the pure acoustic method increases when the SNR is less than 10dB, for example, it can reach 6cm to 8cm; scene adaptability defects. The lack of a multimodal data verification mechanism cannot resolve positioning conflicts in scenarios such as sound source movement and multiple people speaking at the same time.
[0004] Therefore, it is necessary to provide a new sound source localization method and system based on fusion confidence to solve the above problems. Summary of the Invention
[0005] The present invention provides a sound source localization method and system based on fusion confidence to solve technical problems in the prior art such as large positioning errors caused by the singleness of magic fetus data, and the inability to accurately locate scenes such as sound source movement and multiple people speaking at the same time due to the lack of a multimodal data verification mechanism. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0006] In a first aspect, the present invention proposes a sound source localization method based on fusion confidence, comprising: when a preset wake-up word is detected, starting sound source localization processing, specifically including calculating the horizontal azimuth angle and determining the sound source localization area based on the detected current voice signal; performing multi-dimensional voice signal extraction on the current voice signal to generate a voiceprint feature vector of a specified dimension, calculating the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; calculating the user preference fusion coefficient corresponding to the current voice signal, and further calculating the fusion confidence of the current voice signal based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; dynamically adjusting the scanning range of the pan-tilt camera based on the calculated fusion confidence of the current voice signal to capture image frames of the current user at fixed intervals during the scanning process, performing face detection, and obtaining a face image to be processed; performing facial feature extraction on the face image to be processed to perform visual identity collaborative confirmation, and when the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, stopping the rotation of the pan-tilt camera and locking the current orientation as the determined sound source position.
[0007] The second aspect of the present invention proposes a sound source localization system based on fusion confidence, which executes the sound source localization method described in the first aspect of the present invention, and the sound source localization system includes: a pan-tilt camera, a first microphone and a second microphone located on both sides of the pan-tilt camera; a sound source localization module, which is used to start the sound source localization processing when a preset wake-up word is detected, specifically including calculating the horizontal azimuth angle and determining the sound source localization area according to the detected current voice signal; a voiceprint recognition module, which is used to perform multi-dimensional voice signal extraction on the current voice signal to generate a voiceprint feature vector of a specified dimension, calculate the existing user voiceprint that matches the current voice signal, and confirm the current user corresponding to the current voice signal; a control processing unit, which is used to calculate the corresponding voice signal corresponding to the current voice signal. The user preference fusion coefficient is calculated, and the fusion confidence of the current voice signal is further calculated based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal; the pan-tilt mechanical structure unit dynamically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so that the visual processing module captures the image frame of the current user at a fixed interval during the scanning process, performs face detection, and obtains the face image to be processed; the determination module is used to extract facial features of the face image to be processed for visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the pan-tilt camera stops rotating and locks the current position as the determined sound source position.
[0008] The third aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the sound source localization method based on fusion confidence described in the first aspect of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the sound source localization method based on fusion confidence described in the first aspect of the present invention is implemented.
[0010] The embodiments of the present invention include the following advantages:
[0011] Compared with the existing technology, the present invention constructs a three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification. It provides a positioning space benchmark through preliminary positioning of the sound source orientation, narrows the user range by combining voiceprint feature matching, and introduces a dynamic scanning compensation mechanism. It automatically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence, which can effectively improve positioning accuracy while reducing hardware complexity.
[0012] In addition, a cross-modal cross-validation channel for voiceprint ID and face ID is established, and the spatial perception limitations of dual microphones are compensated through learning from historical orientation data, forming a closed-loop system for mutual correction of multimodal features to improve user experience. The introduction of voiceprint recognition and face recognition technology solves the problems of dual-microphone arrays being unable to achieve three-dimensional positioning, low positioning accuracy, and low accuracy in noisy environments, thereby improving the usability of the sound source positioning function and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flowchart of an example of a sound source localization method based on fusion confidence of the present invention;
[0014] Figure 2 1 is a schematic diagram of a framework of a specific application example of the sound source localization method based on fusion confidence of the present invention;
[0015] Figure 3 is an example diagram of a sound source localization area determined in the sound source localization method based on fusion confidence of the present invention;
[0016] Figure 4 It is a structural block diagram of the sound source localization system based on fusion confidence of the present invention;
[0017] Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of the present invention;
[0018] Figure 6 is a schematic structural diagram of an embodiment of a computer-readable medium according to the present invention. DETAILED DESCRIPTION
[0019] In the description of specific embodiments, the features, structures, characteristics, or other details of the present invention are described to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from practicing the technical solutions of the present invention without one or more of the specific features, structures, characteristics, or other details.
[0020] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0021] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. Specifically, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. Furthermore, in the accompanying drawings, the horizontal direction is defined as left-to-right, and the vertical direction (or a direction perpendicular to the horizontal direction) is defined as up-to-down.
[0022] It should be understood that while the terms "first," "second," and "third" may be used herein to describe various devices, elements, components, or parts, this should not be construed as limiting. These terms are used to distinguish one from another. For example, a first device could also be referred to as a second device without departing from the essential technical solution of the present invention.
[0023] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0024] In view of the above problems, the present invention proposes a sound source localization method based on fusion confidence. In this method, when a preset wake-up word is detected, the sound source localization process is started, which specifically includes calculating the horizontal azimuth angle and determining the sound source localization area according to the detected current voice signal; performing multi-dimensional voice signal extraction on the current voice signal to generate a specified dimensional voiceprint feature vector, calculating the existing user voiceprint matching the current voice signal to confirm the current user corresponding to the current voice signal; calculating the user preference fusion coefficient corresponding to the current voice signal, and according to the calculated user preference fusion coefficient matching the current voice signal, determining the sound source localization area; performing multi-dimensional voice signal extraction on the current voice signal to generate a specified dimensional voiceprint feature vector, calculating the existing user voiceprint matching the current voice signal to confirm the current user corresponding to the current voice signal; calculating the user preference fusion coefficient corresponding to the current voice signal, and determining the user preference fusion coefficient corresponding to the current voice signal according to the calculated user preference fusion coefficient matching the current voice signal. The matching degree obtained by matching the existing user's voiceprint is used to further calculate the fusion confidence of the current voice signal; according to the calculated fusion confidence of the current voice signal, the scanning range of the pan-tilt camera is dynamically adjusted to capture the image frame of the current user at fixed intervals during the scanning process, perform face detection, and obtain the face image to be processed; facial features are extracted from the face image to be processed to perform visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the rotation of the pan-tilt camera is stopped and the current orientation is locked as the determined sound source position. By constructing a three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification, the positioning space benchmark is provided through preliminary positioning of the sound source orientation, the user range is narrowed by combining voiceprint feature matching, and a dynamic scanning compensation mechanism is introduced. The scanning range of the pan-tilt camera is automatically adjusted according to the calculated fusion confidence, which can effectively improve the positioning accuracy while reducing the hardware complexity.
[0025] Example 1
[0026] Refer to the following Figure 1 、 Figure 2 、 Figure 3 , the contents of the present invention will be described in detail.
[0027] Figure 1 This is a flowchart of an example of a sound source localization method based on fusion confidence of the present invention. Figure 2 It is a schematic diagram of a framework of a specific application example of the sound source localization method of the present invention. Figure 3 This is an example diagram of the sound source localization area determined in the sound source localization method based on fusion confidence of the present invention.
[0028] Reference Figure 1 and Figure 2 In step S101, when a preset wake-up word is detected, the sound source localization process is started, which specifically includes calculating the horizontal azimuth angle and determining the sound source localization area based on the detected current voice signal.
[0029] exist Figure 2In an application example, the invention includes a pan-tilt camera, a first microphone and a second microphone located on both sides of the pan-tilt camera. The first microphone and the second microphone are symmetrically arranged relative to the pan-tilt camera to form a dual-microphone array.
[0030] For voice devices in smart home scenarios, when a user says a preset wake-up word (such as "Hello **"), the sound source localization module, voiceprint recognition module, and visual processing module (with integrated facial recognition technology) work together to identify the sound source location of the current voice signal (i.e., the current user's location). The PTZ camera is then automatically controlled to precisely steer toward the user, allowing the user to naturally communicate face-to-face with the voice device without having to move or search for the device, enhancing the convenience and experience of interaction. The voice devices mentioned above include speakers, TV boxes, and PTZ cameras.
[0031] Specifically, when a preset wake-up word (such as "Hello **") is detected, the sound source localization processing is started, which specifically includes calculating the horizontal azimuth angle and determining the sound source localization area based on the detected voice signal.
[0032] The horizontal azimuth angle of the current voice signal is calculated using the following expression:
[0033] θ = arcsin(c * Δt / d) (1)
[0034] Among them, θ represents the angle formed by the sound source direction of the current voice signal and the central axis of the pan-tilt camera perpendicular to the horizontal direction, also known as the horizontal azimuth, θ∈[-90°,90°]; c represents the speed of sound propagation, in meters per second; d represents the distance between the first microphone and the second microphone in the horizontal direction, in meters, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal.
[0035] In this example, the value of d is preferably 8 meters. Preferably, the horizontal azimuth angle θ∈[-75°, 75°].
[0036] It should be noted that, through historical detection data, it is determined that when the horizontal azimuth angle θ∈[-75°,75°], the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal can be measured more accurately, and the calculation error is small, which can also provide a more reliable data basis for subsequent voiceprint feature extraction and fusion confidence assessment.
[0037] According to the calculated horizontal azimuth angle, the sound source localization area is determined with the central axis of the pan-tilt camera (such as Figure 3The Z axis is represented by a dotted line, and the center of the pan-tilt camera is point O) and the horizontal azimuth angle area is formed near the first microphone, and the horizontal azimuth angle area (for example, "θ") is formed near the center axis of the pan-tilt camera and the second microphone as the sound source localization area. For details, see Figure 3 The "OHL region" is shown.
[0038] When the signal-to-noise ratio of the current speech signal is greater than or equal to a set threshold, the sound source localization score of the current speech signal is greater than or equal to 0.8.
[0039] Optionally, the threshold is set to be 20 dB to 25 dB. In this example, the threshold is set to be 20 dB.
[0040] It should be noted that the above descriptions are provided as optional examples and should not be construed as limiting the present invention.
[0041] Next, in step S102, multi-dimensional voice signal extraction is performed on the current voice signal to generate a voiceprint feature vector of a specified dimension, and an existing user voiceprint matching the current voice signal is calculated to confirm the current user corresponding to the current voice signal.
[0042] Mel-frequency cepstral coefficients are used to extract specific dimensional speech features of the current speech signal, and a multidimensional voiceprint feature vector is generated through a pre-trained model. The dimension of the multidimensional voiceprint feature vector is greater than the dimension of the specific dimensional speech feature.
[0043] In a specific embodiment, Mel-frequency cepstral coefficients (MFCC) are used to extract, for example, 39-dimensional speech features, i.e., specific dimensional speech features, of the current speech signal. The 39-dimensional speech features include static spectrum information and dynamic change information, specifically twelve cepstral coefficients, energy, first-order difference, and second-order difference corresponding to the current speech signal.
[0044] Based on a CNN network, a pre-trained model is trained using a training dataset to generate a multi-dimensional voiceprint feature vector corresponding to the current voice signal. The pre-trained model comprises an input layer, a five-layer convolutional layer, a three-layer fully connected layer, and an output layer. The training dataset includes voice data (including voice signals) annotated with a user's voiceprint identifier (e.g., user voiceprint ID). In the multi-classification task of model training, voice data annotated with the target user's identifier (e.g., user voiceprint ID) serves as positive samples, while voice data from all other users serves as negative samples.
[0045] Specifically, the current voice signal and the specific dimensional voice features corresponding to the current voice signal are input into the pre-trained model to obtain a multi-dimensional voiceprint feature vector corresponding to the current voice signal. The multi-dimensional voiceprint feature vector is, for example, a 128-dimensional voiceprint feature vector.
[0046] Furthermore, the calculated multi-dimensional voiceprint feature vector is used to perform similarity matching calculation with existing user voiceprints in the feature database, specifically calculating the matching degree of the existing user voiceprint that matches the current voice signal to determine the matching existing user voiceprint.
[0047] For similarity matching calculations, when the calculated matching degree is greater than the matching degree threshold (e.g., greater than or equal to 0.85), an existing user voiceprint that matches the current voice signal is determined. When the calculated matching degree is less than the matching degree threshold (e.g., less than 0.85), the next existing user voiceprint is queried from the feature database and the similarity matching calculations described above are performed until an existing user voiceprint that matches the current voice signal is determined.
[0048] For the feature database, each user is guided to read the wake-up word aloud to form a voice signal, and the voice device collects, for example, an audio clip with a signal-to-noise ratio greater than or equal to 25 dB, and extracts the Mel-frequency cepstrum features of the formed voice signal. The pre-trained model is used to generate a voiceprint feature vector corresponding to the formed voice signal. Each voiceprint feature vector is bound to the corresponding user account, binding information is generated, and the binding information is stored to establish a feature database. Alternatively, by recording the user voiceprint of each user, extracting the user voiceprint identifier of each user, forming binding information between the user voiceprint identifier and the user account, and storing the binding information to establish a feature database.
[0049] For example, the feature database includes user information, user account information (such as user account, user ID, etc.), one or more voiceprint feature vectors corresponding to each user account, and the like.
[0050] It should be noted that the above descriptions are provided as optional examples and should not be construed as limiting the present invention.
[0051] Next, in step S103, a user preference fusion coefficient corresponding to the current voice signal is calculated, and the fusion confidence of the current voice signal is further calculated based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal.
[0052] Based on the user's orientation preference data in the historical interaction data between the user and the home device (for example, the average value of the commonly used horizontal orientation angles in the historical time period of one month, two months, three months, six months, or twelve months from the current time), and the proportion of the user's actual default events in the total interaction samples, i.e., the default rate, an orientation correction coefficient is generated. Specifically, the following expression is used to calculate the orientation correction coefficient:
[0053] α=q×(1-δ / δ max )(2)
[0054] Wherein, α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angle in a specific time period, α∈[0.05,0.2]; q represents the optimal coefficient, which is determined by regression analysis and fitting of delay data and / or intensity data at different angles; δ max It represents the historical maximum value of the proportion of actual default events of users counted in all interaction samples, that is, the historical highest default rate; δ represents the default rate of users whose current user wakes up the voice device outside the specified location.
[0055] It should be noted that, in the present invention, the value of q in the above expression (2) is preferably 0.15, and the optimal coefficient q is obtained by calculating based on data regression analysis. Specifically, 200 sets of delay data and / or intensity data at different angles of home scenes are collected, and the optimal coefficient obtained by regression analysis fitting is used to minimize the positioning error. In addition, the actual user default event refers to the position where the user wakes up the voice device is not within the specified position (for example, not within the angle range formed by the specified horizontal angle, the specified horizontal angle is within the range of -75° to 75°, optionally, within the range of -77° to 73°. When it is not within the range formed by the above-mentioned specified horizontal angle, it is an actual user default event. The above is only explained as an optional example and cannot be understood as a limitation of the present invention.
[0056] Next, the fusion confidence of the current voice signal is further calculated based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal, the signal-to-noise ratio determined at the time of the current voice signal, and the sound source localization score generated by the time difference, including:
[0057] The fusion confidence of the current speech signal is calculated using the following expression:
[0058] C fusion =k*S loc +w*S voice +α*Δθ (3)
[0059] Among them, C fusionCharacterizes the fusion confidence of the current speech signal; S loc S is a sound source localization score generated by the signal-to-noise ratio and time difference determined when the current voice signal is collected by the first microphone and the second microphone; voice represents the calculated matching degree between the current voice signal and the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth and the mean value of the historical horizontal azimuth within a specific time period, Δθ=|θ-μ|, θ represents the current horizontal azimuth, μ represents the mean value of the historical horizontal azimuth distribution within a specific time period, that is, the mean value of the historical horizontal azimuth; k represents the first parameter, specifically the first parameter corresponding to the sound source localization score generated by the signal-to-noise ratio and time difference determined when the current voice signal is collected by dual microphones; w represents the second parameter, specifically the second parameter corresponding to the calculated matching degree between the current voice signal and the existing user voiceprint; α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation value between the current horizontal azimuth and the mean value of the historical horizontal azimuth within a specific time period.
[0060] Optionally, the value range of k is 0.55 to 0.75, the value range of w is 0.25 to 0.45, and the following expression is satisfied: k+w=1.
[0061] Preferably, the value of k is 0.6, and the value range of w is 0.4.
[0062] It should be noted that the value of k is 0.6 and the value of w is 0.4. These are obtained by collecting 200 sets of delay data and / or intensity data from different angles of home scenes and fitting them using regression analysis to obtain the optimal coefficients to minimize positioning error. The preferred values of k and w correspond to the optimal coefficients and are determined based on the optimal coefficients and the minimum positioning error. The above is provided as an optional example and is not to be construed as a limitation of the present invention.
[0063] Next, in step S104, the scanning range of the pan-tilt camera is dynamically adjusted according to the calculated fusion confidence of the current voice signal to capture image frames of the current user at fixed intervals during the scanning process, perform face detection, and obtain the face image to be processed.
[0064] Specifically, the scanning range of the pan-tilt camera is dynamically adjusted according to the calculated fusion confidence of the current speech signal.
[0065] When the calculated fusion confidence of the current speech signal is greater than or equal to 0.9, the scanning angle of the pan-tilt camera is increased by 5° or decreased by 5°.
[0066] When the calculated fusion confidence of the current speech signal is greater than or equal to 0.7 and less than 0.9, the scanning angle of the pan-tilt camera is increased by 10° or decreased by 10°.
[0067] When the calculated fusion confidence of the current speech signal is less than 0.7, the scanning angle of the pan-tilt camera is increased or decreased by 15°.
[0068] Furthermore, during the scanning process of the pan-tilt camera, image frames of the current user are captured at fixed intervals, and face detection is performed to obtain a face image to be processed.
[0069] The fixed interval is, for example, 0.1s to 0.5s.
[0070] It should be noted that the above description is merely provided as an optional example and should not be construed as a limitation to the present invention.
[0071] Next, in step S105, facial features are extracted from the face image to be processed to perform visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the rotation of the pan-tilt camera is stopped and the current orientation is locked, that is, the sound source position is determined.
[0072] Specifically, facial features are extracted from the face image to be processed to perform visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the rotation of the pan-tilt camera is stopped and the current orientation is locked as the determined sound source position.
[0073] Specifically, facial features are extracted from the processed face image, and a FaceNet model is used to extract, for example, a 512-dimensional feature vector. The 512-dimensional feature vector includes geometric features such as eye distance and nose tip coordinates, texture features (local binary pattern LBP) and global features (facial contour encoding).
[0074] For visual identity collaborative confirmation, cross-modal verification is performed to specifically compare whether the user identifier (e.g., UserID) associated with the facial features is consistent with the user identifier (e.g., UserID) corresponding to the voiceprint feature vector of the current voice information. Specifically, the consistency threshold (e.g., 0.92) is used, or the Euclidean distance is used to judge the consistency (e.g., by judging when the calculated Euclidean distance is less than or equal to 1.2) to determine whether the user identifier (e.g., UserID) associated with the facial features is consistent with the user identifier corresponding to the voiceprint feature vector of the current voice information.
[0075] In one specific embodiment, the Euclidean distance between the user identifier corresponding to the human face image to be processed and the user identifier corresponding to the voiceprint feature vector of the current voice information is calculated. When the calculated Euclidean distance is less than or equal to 1.2, it indicates that the visual identity collaborative confirmation result of the facial image to be processed meets the identity consistency condition. If the calculated Euclidean distance is greater than 1.2, it indicates that the visual identity collaborative confirmation result of the facial image to be processed does not meet the identity consistency condition.
[0076] In another specific embodiment, a similarity is calculated between a user identifier (e.g., UserID) associated with the facial features and the user identifier (e.g., UserID) corresponding to the voiceprint feature vector of the current voice information. When the calculated similarity is greater than a consistency threshold (e.g., 0.92), it indicates that the visual identity collaborative confirmation result of the facial image to be processed meets the identity consistency condition. When the calculated similarity is less than or equal to the consistency threshold (e.g., 0.92), it indicates that the visual identity collaborative confirmation result of the facial image to be processed does not meet the identity consistency condition.
[0077] When the identity consistency condition is met, the control processing unit controls the pan-tilt mechanical structure unit to immediately stop the horizontal rotation of the pan-tilt camera and lock the current position, specifically facing the center of the camera towards the current user, thereby using the locked current position as the determined sound source position.
[0078] It should be noted that the consistency threshold is determined through test set optimization of the cross-modal validation model. Specifically, while maintaining a 98% true acceptance rate (TAR), the false acceptance rate (FAR) is reduced to less than or equal to 0.1%. Under these conditions, the consistency threshold is preferably 0.92. The above is provided as an optional example only and should not be construed as limiting the present invention.
[0079] Compared with the existing technology, the present invention constructs a three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification. It provides a positioning space benchmark through preliminary positioning of the sound source orientation, narrows the user range by combining voiceprint feature matching, and introduces a dynamic scanning compensation mechanism. It automatically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence, which can effectively improve positioning accuracy while reducing hardware complexity.
[0080] In addition, a cross-modal cross-validation channel for voiceprint ID and face ID is established, and the spatial perception limitations of dual microphones are compensated through learning from historical orientation data, forming a closed-loop system for mutual correction of multimodal features to improve user experience. The introduction of voiceprint recognition and face recognition technology solves the problems of dual-microphone arrays being unable to achieve three-dimensional positioning, low positioning accuracy, and low accuracy in noisy environments, thereby improving the usability of the sound source positioning function and user experience.
[0081] Example 2
[0082] The following are system embodiments of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the system embodiments of the present invention, please refer to the method embodiments of the present invention.
[0083] Figure 4 FIG. 1 is a structural diagram of an example of a sound source localization system based on fusion confidence according to the present invention. Figure 4 , a sound source localization system 400 based on fusion confidence is described, and the sound source localization system 400 executes the sound source localization method described in Example 1 of the present invention.
[0084] Specifically, the sound source localization system 400 includes a pan-tilt camera 410, a first microphone 420 and a second microphone 430 located on both sides of the pan-tilt camera, a sound source localization module 440, a voiceprint recognition module 450, a control processing unit 460, a pan-tilt mechanical structure unit 470 and a determination module 480.
[0085] Furthermore, the sound source localization module 440 is used to start the sound source localization processing when a preset wake-up word is detected, specifically including calculating the horizontal azimuth angle and determining the sound source localization area based on the detected current voice signal. The voiceprint recognition module 450 is used to perform multi-dimensional voice signal extraction on the current voice signal to generate a voiceprint feature vector of a specified dimension, and calculate the existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal. The control processing unit 460 is used to calculate the user preference fusion coefficient corresponding to the current voice signal, and further calculate the fusion confidence of the current voice signal based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal. The pan-tilt mechanical structure unit 470 dynamically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal, so that the visual processing module captures the image frame of the current user at a fixed interval during the scanning process, performs face detection, and obtains the face image to be processed. The determination module 480 is used to extract facial features of the face image to be processed for visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the rotation of the pan-tilt camera is stopped and the current position is locked as the determined sound source position.
[0086] Optionally, the sound source localization system is used in conjunction with an audio device, and specifically the sound source localization module is used to monitor a preset wake-up word used by the current user to wake up the audio device, wherein the audio device includes an audio device and a TV box.
[0087] It should be noted that in other embodiments, the audio device is an integrated product, specifically including an audio device, a TV box, and a camera (such as a pan-tilt camera). The above is only provided as an optional example and should not be construed as limiting the present invention.
[0088] According to an optional implementation manner, the fusion confidence of the current voice signal is further calculated based on the matching degree obtained by calculating the existing user voiceprint that matches the current voice signal.
[0089] The fusion confidence of the current speech signal is calculated using the following expression:
[0090] C fusion =k*S loc +w*S voice +α*Δθ
[0091] Among them, C fusion Characterizes the fusion confidence of the current speech signal; S loc S is a sound source localization score generated by the signal-to-noise ratio and time difference determined when the current voice signal is collected by the first microphone and the second microphone; voice represents the calculated matching degree between the current voice signal and the existing user voiceprint; Δθ represents the absolute deviation value between the current horizontal azimuth and the mean value of the historical horizontal azimuth within a specific time period, Δθ=|θ-μ|, θ represents the current horizontal azimuth, μ represents the mean value of the historical horizontal azimuth distribution within a specific time period, that is, the mean value of the historical horizontal azimuth; k represents the first parameter, specifically the first parameter corresponding to the sound source localization score generated by the signal-to-noise ratio and time difference determined when the current voice signal is collected by dual microphones; w represents the second parameter, specifically the second parameter corresponding to the calculated matching degree between the current voice signal and the existing user voiceprint; α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation value between the current horizontal azimuth and the mean value of the historical horizontal azimuth within a specific time period.
[0092] According to an optional embodiment, a location correction coefficient is generated based on the user's location preference data in the historical interaction data between the user and the home device, and the proportion of the user's actual default events in all interaction samples, that is, the default rate. Specifically, the location correction coefficient is calculated using the following expression:
[0093] α=q×(1-δ / δ max )
[0094] Wherein, α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angle in a specific time period, α∈[0.05,0.2]; q represents the optimal coefficient, which is determined by regression analysis and fitting of delay data and / or intensity data at different angles; δ max It represents the historical maximum value of the proportion of actual default events of users counted in all interaction samples, that is, the historical highest default rate; δ represents the default rate of users whose current user wakes up the voice device outside the specified location.
[0095] The value range of k is 0.55 to 0.75, the value range of w is 0.25 to 0.45, and the following expression is satisfied: k+w=1.
[0096] The horizontal azimuth angle of the current voice signal is calculated using the following expression:
[0097] θ = arcsin(c * Δt / d)
[0098] Among them, θ represents the angle formed by the sound source direction of the current voice signal and the central axis of the pan-tilt camera perpendicular to the horizontal direction, θ∈[-90°,90°]; c represents the speed of sound propagation, in meters per second; d represents the horizontal distance between the first microphone and the second microphone, in meters, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal.
[0099] According to the calculated horizontal azimuth angle, the sound source positioning area is determined, and the area range of the horizontal azimuth angle formed by the central axis of the pan-tilt camera close to the first microphone and the area range of the horizontal azimuth angle formed by the central axis of the pan-tilt camera close to the second microphone are used as the determined sound source position.
[0100] According to an optional implementation manner, when the signal-to-noise ratio of the current speech signal is greater than or equal to a set threshold, the sound source localization score of the current speech signal is greater than or equal to 0.8, and the set threshold is 20 dB to 25 dB.
[0101] According to an optional implementation, the scanning range of the pan-tilt camera is dynamically adjusted based on the calculated fusion confidence of the current speech signal.
[0102] When the calculated fusion confidence of the current speech signal is greater than or equal to 0.9, the scanning angle of the pan-tilt camera is increased by 5° or decreased by 5°.
[0103] When the calculated fusion confidence of the current speech signal is greater than or equal to 0.7 and less than 0.9, the scanning angle of the pan-tilt camera is increased by 10° or decreased by 10°.
[0104] When the calculated fusion confidence of the current speech signal is less than 0.7, the scanning angle of the pan-tilt camera is increased or decreased by 15°.
[0105] Mel-frequency cepstral coefficients are used to extract specific dimensional speech features of the current speech signal, and a multidimensional voiceprint feature vector is generated through a pre-trained model. The dimension of the multidimensional voiceprint feature vector is greater than the dimension of the specific dimensional speech feature.
[0106] Preferably,
[0107] It should be noted that due to Figure 4 The sound source localization method performed by the sound source localization system based on fusion confidence is the same as Figure 1 The sound source localization methods in the examples are substantially the same, and therefore, descriptions of the same parts are omitted.
[0108] Compared with the existing technology, the present invention constructs a three-level collaborative architecture of acoustic positioning, facial feature recognition and visual verification. It provides a positioning space benchmark through preliminary positioning of the sound source orientation, narrows the user range by combining voiceprint feature matching, and introduces a dynamic scanning compensation mechanism. It automatically adjusts the scanning range of the pan-tilt camera according to the calculated fusion confidence, which can effectively improve positioning accuracy while reducing hardware complexity.
[0109] In addition, a cross-modal cross-validation channel for voiceprint ID and face ID is established, and the spatial perception limitations of dual microphones are compensated through learning from historical orientation data, forming a closed-loop system for mutual correction of multimodal features to improve user experience. The introduction of voiceprint recognition and face recognition technology solves the problems of dual-microphone arrays being unable to achieve three-dimensional positioning, low positioning accuracy, and low accuracy in noisy environments, thereby improving the usability of the sound source positioning function and user experience.
[0110] Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of the present invention.
[0111] like Figure 5 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.
[0112] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.
[0113] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).
[0114] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0115] It should be understood that Figure 5 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.
[0116] Through the above description of the embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Figure 6 As shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.
[0117] The software product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0118] The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal propagated may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0119] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0120] The computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the computer-readable medium implements the data interaction method of the present disclosure.
[0121] Those skilled in the art will appreciate that the modules described above can be distributed in the device according to the description of the embodiment, or can be modified accordingly to be used in one or more devices that are different from the embodiment. The modules of the above embodiment can be combined into one module or further divided into multiple submodules.
[0122] From the above description of the embodiments, those skilled in the art will readily appreciate that the exemplary embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored on a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes commands that cause a computing device (such as a personal computer, server, mobile terminal, or network device) to execute the methods according to the embodiments of the present invention.
[0123] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs.
[0124] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless the context dictates otherwise. The illustrated embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein.
[0125] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A sound source localization method based on fusion confidence, characterized in that: include: When a preset wake-up word is detected, the sound source localization process is started, specifically including calculating the horizontal azimuth angle and determining the sound source localization area based on the detected current voice signal; Perform multi-dimensional voice signal extraction on the current voice signal to generate a voiceprint feature vector of a specified dimension, and calculate an existing user voiceprint that matches the current voice signal to confirm the current user corresponding to the current voice signal; Calculating an orientation correction coefficient corresponding to the current voice signal, and further calculating a fusion confidence of the current voice signal based on a matching degree obtained by calculating an existing user voiceprint that matches the current voice signal, including: The fusion confidence of the current speech signal is calculated using the following expression: C fusion =k*S loc +w*S vioice +a*Δθ Among them, C fusion Characterizes the fusion confidence of the current speech signal; S loc S is a sound source localization score generated by the signal-to-noise ratio and time difference determined when the current voice signal is collected by the first microphone and the second microphone; vioice represents the calculated matching degree between the current voice signal and the existing user voiceprint; Δθ represents the absolute deviation between the current horizontal azimuth angle and the mean of the historical horizontal azimuth angles within a specific time period, Δθ=|θ-μ|, θ represents the current horizontal azimuth angle, μ represents the mean of the historical horizontal azimuth angle distribution within a specific time period, that is, the mean of the historical horizontal azimuth angles; k represents a first parameter corresponding to the sound source localization score generated by the signal-to-noise ratio and time difference determined when collecting the current voice signal through dual microphones; w represents a second parameter corresponding to the calculated matching degree between the current voice signal and the existing user voiceprint; α represents an azimuth correction coefficient, which is used to correct the absolute deviation between the current horizontal azimuth angle and the mean of the historical horizontal azimuth angles within a specific time period; Based on the calculated fusion confidence of the current voice signal, the scanning range of the pan-tilt camera is dynamically adjusted to capture image frames of the current user at fixed intervals during the scanning process, perform face detection, and obtain a face image to be processed; Facial features are extracted from the face image to be processed to perform visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed satisfies the identity consistency condition, the rotation of the pan-tilt camera is stopped and the current orientation is locked as the determined sound source position.
2. The sound source localization method according to claim 1, wherein: Further including: Based on the user's location preference data in the historical interaction data between the user and the home device, and the proportion of the user's actual default events in all interaction samples, i.e., the default rate, a location correction coefficient is generated. Specifically, the following expression is used to calculate the location correction coefficient: α=q×(1-δ / δ max ) Wherein, α represents the azimuth correction coefficient, which is specifically used to correct the absolute deviation between the current horizontal azimuth angle and the mean value of the historical horizontal azimuth angle in a specific time period, α∈[0.05,0.2]; q represents the optimal coefficient, which is determined by regression analysis and fitting of delay data and / or intensity data at different angles; δ max It represents the historical maximum value of the proportion of actual default events of users counted in all interaction samples, that is, the historical highest default rate; δ represents the default rate of users whose current user wakes up the voice device outside the specified location.
3. The sound source localization method according to claim 1, wherein: Further including: The value range of k is 0.55 to 0.75, the value range of w is 0.25 to 0.45, and the following expression is satisfied: k+w=1.
4. The sound source localization method according to claim 1, wherein: Further including: The horizontal azimuth angle of the current voice signal is calculated using the following expression: θ = arcsin(c * Δt / d) Wherein, θ represents the angle formed by the sound source direction of the current voice signal and the central axis of the pan-tilt camera perpendicular to the horizontal direction, θ∈[-90°,90°]; c represents the speed of sound propagation, in meters per second; d represents the horizontal distance between the first microphone and the second microphone, in meters, and the value range of d is 8 meters to 15 meters; Δt represents the time difference between the time when the first microphone collects the current voice signal and the time when the second microphone collects the current voice signal; Based on the calculated horizontal azimuth angle, the sound source localization area is determined, and the area range with the horizontal azimuth angle formed by the central axis of the pan-tilt camera close to the first microphone and the area range with the horizontal azimuth angle formed by the central axis of the pan-tilt camera close to the second microphone are used as the determined sound source localization area.
5. The sound source localization method according to claim 4, characterized in that: Further including: When the signal-to-noise ratio of the current speech signal is greater than or equal to a set threshold, the sound source localization score of the current speech signal is greater than or equal to 0.8, and the set threshold is 20 dB to 25 dB.
6. The sound source localization method according to claim 1, wherein: The dynamically adjusting the scanning range of the pan-tilt camera according to the calculated fusion confidence of the current voice signal includes: When the calculated fusion confidence of the current speech signal is greater than or equal to 0.9, the scanning angle of the pan-tilt camera is increased or decreased by 5°; When the calculated fusion confidence of the current speech signal is greater than or equal to 0.7 and less than or equal to 0.9, the scanning angle of the pan-tilt camera is increased or decreased by 10°; When the calculated fusion confidence of the current speech signal is greater than or equal to 0.7, the scanning angle of the pan-tilt camera is increased or decreased by 15°.
7. The sound source localization method according to claim 1, wherein: Further including: Mel-frequency cepstral coefficients are used to extract specific dimensional speech features of the current speech signal, and a multidimensional voiceprint feature vector is generated through a pre-trained model. The dimension of the multidimensional voiceprint feature vector is greater than the dimension of the specific dimensional speech feature.
8. A sound source localization system based on fusion confidence, characterized in that: It implements the sound source localization method according to any one of claims 1 to 7, and the sound source localization system comprises: a pan-tilt camera, and a first microphone and a second microphone located on both sides of the pan-tilt camera; The sound source localization module is used to start the sound source localization processing when the preset wake-up word is detected. Specifically, it calculates the horizontal azimuth angle and determines the sound source localization area based on the detected current voice signal; A voiceprint recognition module is used to perform multi-dimensional voice signal extraction on the current voice signal to generate a voiceprint feature vector of a specified dimension, calculate the existing user voiceprint that matches the current voice signal, and confirm the current user corresponding to the current voice signal; a control processing unit, configured to calculate an orientation correction coefficient corresponding to the current voice signal, and further calculate a fusion confidence of the current voice signal based on a matching degree obtained by calculating an existing user voiceprint that matches the current voice signal; The pan-tilt mechanical structure unit dynamically adjusts the scanning range of the pan-tilt camera based on the calculated fusion confidence of the current voice signal, so that the visual processing module captures image frames of the current user at fixed intervals during the scanning process, performs face detection, and obtains a face image to be processed; The determination module is used to extract facial features of the face image to be processed to perform visual identity collaborative confirmation. When the visual identity collaborative confirmation result of the face image to be processed meets the identity consistency condition, the pan-tilt camera is stopped from rotating and the current orientation is locked as the determined sound source position.
Citation Information
Patent Citations
Intelligent robot rotation method based on sound source positioning and face detection
CN106292732A
Object recognition method and device, storage medium and terminal
CN108305615A
Sound source positioning method and audio equipment
CN113640744A
Bimodal identity authentication method and device and storage medium
CN114398611A