Voiceprint recognition system and method

Through the combination of optical resonant acoustic sensor and data processing unit, multiple resonant modes and differential detection technologies are used to generate voiceprint QR codes, solving the problem of low accuracy of sound signal source recognition in low signal-to-noise ratio environments, and achieving high-precision weak sound signal recognition.

CN120412592BActive Publication Date: 2025-09-02SHANGHAI JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510912605.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-02
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The prior art has low accuracy in identifying sound signal sources in low signal-to-noise ratio environments, mainly due to the poor quality of the collected sound signals.

Method used

The optical resonant acoustic sensor uses multiple resonant modes to detect the frequency and difference detection through two reference lights with orthogonal phases to generate a voiceprint QR code, and match the target voiceprint QR code from the voiceprint QR code database to determine the source of the sound signal.

Benefits of technology

Under low signal-to-noise ratio conditions, high-precision identification of weak sound signals is achieved, the accuracy of identification of sound signal sources is improved, noise is suppressed at the detection end through physical means, and signal quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412592B_ABST
    Figure CN120412592B_ABST
Patent Text Reader

Abstract

The present invention provides a voiceprint recognition system and method to solve the technical problem of low accuracy in identifying the source of sound signals in the existing technology, and can be applied to the field of sound signal sensing and recognition technology. The system includes: an optical resonant sound sensor uses multiple detection lights in multiple resonant modes to detect sound signals from a target scene to obtain sound modulated signal light; a modulated signal detection unit beats the sound modulated signal light with two phase-orthogonal reference lights to obtain two beat signals, and performs differential detection on the obtained two beat signals to obtain a first detection result and a second detection result; a data processing unit generates a voiceprint QR code based on the first detection result and the second detection result in a preset time period; a target voiceprint QR code that matches the voiceprint QR code is determined from multiple preset voiceprint QR codes included in a voiceprint QR code database; and an object that matches the sound signal is determined based on the target voiceprint QR code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of acoustic signal sensing and recognition, and more specifically, to a voiceprint recognition system and method. Background Art

[0002] Sound signals are crucial for human exploration of nature and acquisition of information, and are widely used in scientific research, daily life, industrial production, healthcare, and national defense. In many practical applications, the sound signal to be measured may be weak and accompanied by background noise. Therefore, it is often necessary to extract weak signals from environments with low signal-to-noise ratios in order to further determine their source.

[0003] Existing technologies for identifying sound signals typically rely on post-processing methods based on detection data, such as feature extraction and machine learning, to identify the source of sound signals. However, existing technologies rely on the quality of detection data, which severely impacts the accuracy of source identification, resulting in low accuracy. Summary of the Invention

[0004] In view of this, the present invention provides a voiceprint recognition system and method.

[0005] According to one aspect of the present invention, a voiceprint recognition system is provided, comprising: an optical resonant acoustic sensor for detecting sound signals from a target scene using multiple detection lights in multiple resonant modes to obtain sound modulated signal light; a modulated signal detection unit for beating the sound modulated signal light with two phase-orthogonal reference lights to obtain two beat signals, and performing differential detection on the obtained two beat signals to obtain a first detection result and a second detection result, wherein the detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval; a data processing unit for generating a voiceprint QR code based on the first detection result and the second detection result in a preset time period; determining a target voiceprint QR code that matches the voiceprint QR code from a plurality of preset voiceprint QR codes included in a voiceprint QR code database; and determining an object that matches the sound signal based on the target voiceprint QR code.

[0006] According to an embodiment of the present invention, the data processing unit generates a voiceprint QR code based on the first detection result and the second detection result of a preset time period, including: dividing the preset time period into multiple sub-time periods; for each of the multiple sub-time periods, performing a fast Fourier transform on the first detection result in each sub-time period to obtain a first transformation result, and performing a fast Fourier transform on the second detection result in each sub-time period to obtain a second transformation result; calculating the square sum of the first transformation result and the second transformation result to obtain a square sum result; generating the voiceprint QR code based on the square sum results in the multiple sub-time periods.

[0007] According to an embodiment of the present invention, the above-mentioned multiple sub-periods include M, where M is an integer greater than 1; the above-mentioned data processing unit calculates the square sum of the above-mentioned first transformation result and the above-mentioned second transformation result to obtain the square sum result, including: in each sub-period, for each detection frequency among the multiple detection frequencies, the square sum of the first transformation result and the second transformation result corresponding to each detection frequency is calculated to obtain the square sum result corresponding to each detection frequency, wherein the above-mentioned multiple detection frequencies correspond one-to-one to the above-mentioned multiple detection lights; the above-mentioned data processing unit generates the above-mentioned voiceprint QR code according to the square sum results in the above-mentioned multiple sub-periods, including: determining the data included in the mth row of the above-mentioned voiceprint QR code according to the square sum results corresponding to the above-mentioned multiple detection frequencies in the mth sub-period, where m is an integer greater than 1 and less than or equal to M.

[0008] According to an embodiment of the present invention, the system further includes: a first optical frequency comb for outputting a plurality of signal lights; a polarization controller for performing polarization modulation on the plurality of signal lights to obtain a plurality of polarization-modulated signal lights; and the optical resonant acoustic sensor is further configured to be excited into the plurality of resonant modes by the plurality of polarization-modulated signal lights and to generate the plurality of detection lights respectively in the plurality of resonant modes.

[0009] According to an embodiment of the present invention, the system further includes: a second optical frequency comb for outputting initial reference light; the modulation signal detection unit is further used to split and phase-shift the initial reference light to obtain the two reference lights.

[0010] According to an embodiment of the present invention, the above-mentioned modulation signal detection unit includes: a first beam splitter, used to split the above-mentioned acoustic modulation signal light into a first signal light and a second signal light; a second beam splitter, used to split the above-mentioned initial reference light into a first initial sub-light and a second initial sub-light; a first phase shifter, used to perform a first phase shift on the above-mentioned first initial sub-light to obtain a first reference light; a second phase shifter, used to perform a second phase shift on the above-mentioned second initial sub-light to obtain a second reference light, wherein the phase of the above-mentioned second reference light is orthogonal to the phase of the above-mentioned first reference light; a third beam splitter, used to beat the above-mentioned first signal light and the above-mentioned first reference light to obtain a first beat signal; and a fourth beam splitter, used to beat the above-mentioned second signal light and the above-mentioned second reference light to obtain a second beat signal.

[0011] According to an embodiment of the present invention, the above-mentioned modulated signal detection unit further includes: a first balanced photodetector, used to perform differential detection on the above-mentioned first beat frequency signal to obtain the above-mentioned first detection result; and a second balanced photodetector, used to perform differential detection on the above-mentioned second beat frequency signal to obtain the above-mentioned second detection result.

[0012] According to an embodiment of the present invention, the optical resonance acoustic sensor includes one of the following: an optical resonance acoustic sensor including a whispering gallery mode microcavity or a photonic crystal microcavity.

[0013] According to an embodiment of the present invention, the data processing unit determines the target voiceprint QR code that matches the voiceprint QR code from the multiple preset voiceprint QR codes included in the voiceprint QR code database, including: calculating the similarities between the voiceprint QR code and the multiple preset voiceprint QR codes respectively to obtain multiple similarities; and determining the target voiceprint QR code from the multiple preset voiceprint QR codes according to the maximum similarity among the multiple similarities.

[0014] According to another aspect of the present invention, a voiceprint recognition method is provided, which is applied to the above-mentioned system, and the method includes: an optical resonant acoustic sensor uses multiple detection lights in multiple resonant modes to detect sound signals from a target scene to obtain sound modulated signal light; a modulated signal detection unit beats the sound modulated signal light with two phase-orthogonal reference lights to obtain two beat signals, and performs differential detection on the obtained two beat signals to obtain a first detection result and a second detection result, wherein the detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval; a data processing unit generates a voiceprint QR code based on the first detection result and the second detection result in a preset time period; determines a target voiceprint QR code that matches the voiceprint QR code from a plurality of preset voiceprint QR codes included in a voiceprint QR code database; and determines an object that matches the sound signal based on the target voiceprint QR code.

[0015] According to an embodiment of the present invention, an optical resonant acoustic sensor utilizes multiple detection lights in multiple resonant modes to detect acoustic signals from a target scene, thereby obtaining acoustically modulated signal light. This allows weak acoustic signal detection and noise suppression to be performed on the target scene at a detection end based on the multiple detection lights of different frequencies. A modulation signal detection unit uses two phase-orthogonal reference lights to beat the acoustically modulated signal light, obtaining two beat signals. The two beat signals are then differentially detected to obtain a first detection result and a second detection result. The detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, which is different from the first predetermined frequency interval. The reference light is used to demodulate the signal lights corresponding to the different detection frequencies in the acoustically modulated signal light, thereby obtaining the first detection result and the second detection result. By utilizing a data processing unit, a voiceprint QR code is generated based on the first detection result and the second detection result of a preset time period. A target voiceprint QR code that matches the voiceprint QR code is determined from a plurality of preset voiceprint QR codes included in a voiceprint QR code database. Based on the target voiceprint QR code, an object that matches the sound signal is determined. The waveform features of a specific sound signal, such as amplitude, frequency, and pauses, can be used to generate a voiceprint QR code. The generated voiceprint QR code is then compared with a pre-calibrated preset voiceprint QR code to determine the object that emits the sound signal. This enables high-precision identification of the source of weak sound signals under low signal-to-noise ratio conditions, thereby improving the accuracy of identifying the source of sound signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other objects, features and advantages of the present invention will become more apparent from the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0017] Figure 1 FIG. 4 shows a structural diagram of a voiceprint recognition system according to an embodiment of the present invention.

[0018] Figure 2 FIG. 4 shows a structural diagram of a voiceprint recognition system according to another embodiment of the present invention.

[0019] Figure 3 A flow chart of a voiceprint recognition method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0020] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0021] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0023] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0024] In the embodiments of the present invention, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken to prevent unauthorized access to user personal information data and maintain the security of user personal information and network security.

[0025] In the embodiment of the present invention, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0026] The quality of collected sound signals significantly impacts the accuracy of existing technologies for identifying the source of sound signals, resulting in low accuracy. Therefore, if the detection end could accurately detect weak sound signals, for example, by leveraging the physical properties of sound signals to identify their source, this could fundamentally address the impact of detection data quality on the accuracy of identifying the source of sound signals.

[0027] Based on the above problems, the present invention uses an optical resonant acoustic sensor to detect weak sound signals, and adopts multiple resonance modes to suppress noise during the detection process, thereby improving the quality of weak sound signals and further improving the accuracy of identifying the source of sound signals.

[0028] Figure 1 FIG. 4 shows a structural diagram of a voiceprint recognition system according to an embodiment of the present invention.

[0029] like Figure 1 As shown, the voiceprint recognition system 100 may include an optical resonance acoustic sensor 110 , a modulated signal detection unit 120 and a data processing unit 130 .

[0030] The optical resonance acoustic sensor 110 can be used to detect sound signals from a target scene using multiple detection lights in multiple resonance modes to obtain acoustically modulated signal light.

[0031] According to an embodiment of the present invention, the optical resonant acoustic sensor 110 is capable of simultaneously exciting multiple resonant modes based on the action of multiple signal lights of different frequencies. The frequencies of the multiple signal lights can be represented in the frequency domain as an optical frequency sequence with equal frequency intervals.

[0032] For example, the plurality of signal lights may be polarization-modulated, so that the plurality of polarization-modulated signal lights can excite a plurality of resonance modes of the optical resonance acoustic sensor 110 .

[0033] According to an embodiment of the present invention, the light field of the signal light It can be expressed as formula (1).

[0034] (1);

[0035] Where j is the optical frequency comb tooth number, is the amplitude of the jth tooth of the signal light optical frequency comb, f s0 is the reference frequency of the signal light optical frequency comb, f sr is the repetition frequency of the signal light optical frequency comb, t is the time, f s0 +jf sr is the frequency of the signal light corresponding to the tooth number j of the optical frequency comb, and i is an imaginary unit.

[0036] According to an embodiment of the present invention, the detection frequency of the detection light corresponding to the optical frequency comb tooth number j is equal to the frequency of the signal light corresponding to the optical frequency comb tooth number j, both being f s0 +jf sr .

[0037] According to the embodiment of the present invention, the sound signal modulates the intensity and phase of the multiple detection lights in multiple resonance modes generated by the optical resonant acoustic sensor 110 to obtain acoustically modulated signal light, whose modulation information is reflected on each comb tooth of the signal light optical frequency comb. It can be expressed as formula (2).

[0038] (2);

[0039] in, and They are the changes in intensity and phase caused by the interaction of the sound signal with each detection light of the optical resonant acoustic sensor.

[0040] The modulation signal detection unit 120 can be used to beat the two phase-orthogonal reference lights with the acoustic modulation signal light respectively to obtain two beat signals, and perform differential detection on the two beat signals respectively to obtain a first detection result and a second detection result, wherein the detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval.

[0041] According to an embodiment of the present invention, the first predetermined frequency interval may be equal to a repetition frequency of the signal light optical frequency comb.

[0042] According to an embodiment of the present invention, the frequency of the modulated signal light corresponding to the optical frequency comb tooth number j is equal to the detection frequency of the detection light corresponding to the optical frequency comb tooth number j. When the detection frequencies of the multiple detection lights are linearly arranged based on the first predetermined frequency interval, the frequencies of the multiple modulated signal sub-lights included in the modulated signal light are linearly arranged based on the first predetermined frequency interval. Since the second predetermined frequency interval is different from the first predetermined frequency interval, the two reference lights with orthogonal phases are respectively beat with the acoustic modulated signal light to obtain two beat signals, and the two beat signals are differentially detected to obtain the first detection result and the second detection result. The first detection result can include the detection results corresponding to each detection frequency, and the second detection result can include the detection results corresponding to each detection frequency.

[0043] The data processing unit 130 can be used to generate a voiceprint QR code based on the first detection result and the second detection result of a preset time period; determine a target voiceprint QR code that matches the voiceprint QR code from a plurality of preset voiceprint QR codes included in the voiceprint QR code database; and determine an object that matches the sound signal based on the target voiceprint QR code.

[0044] According to an embodiment of the present invention, before determining a target voiceprint QR code that matches a voiceprint QR code from a plurality of preset voiceprint QR codes included in a voiceprint QR code database, the sound signal to be identified can be calibrated to generate a voiceprint QR code database so that the voiceprint QR code database can store the association relationship between the target voiceprint QR code and the object corresponding to it.

[0045] According to an embodiment of the present invention, the target voiceprint can be identified in two dimensions by identifying the similarity between the detected voiceprint QR code and multiple preset voiceprint QR codes included in the voiceprint QR code database, thereby identifying the object that matches the sound signal.

[0046] According to an embodiment of the present invention, an optical resonant acoustic sensor utilizes multiple detection lights in multiple resonant modes to detect acoustic signals from a target scene, thereby obtaining acoustically modulated signal light. This allows weak acoustic signal detection and noise suppression to be performed on the target scene at a detection end based on the multiple detection lights of different frequencies. A modulation signal detection unit uses two phase-orthogonal reference lights to beat the acoustically modulated signal light, obtaining two beat signals. The two beat signals are then differentially detected to obtain a first detection result and a second detection result. The detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, which is different from the first predetermined frequency interval. The reference light is used to demodulate the signal lights corresponding to the different detection frequencies in the acoustically modulated signal light, thereby obtaining the first detection result and the second detection result. By utilizing a data processing unit, a voiceprint QR code is generated based on the first detection result and the second detection result of a preset time period. A target voiceprint QR code that matches the voiceprint QR code is determined from a plurality of preset voiceprint QR codes included in a voiceprint QR code database. Based on the target voiceprint QR code, an object that matches the sound signal is determined. The waveform features of a specific sound signal, such as amplitude, frequency, and pauses, can be used to generate a voiceprint QR code. The generated voiceprint QR code is then compared with a pre-calibrated preset voiceprint QR code to determine the object that emits the sound signal. This enables high-precision identification of the source of weak sound signals under low signal-to-noise ratio conditions, thereby improving the accuracy of identifying the source of sound signals.

[0047] According to the embodiments of the present invention, the present invention combines the physical acquisition end and the data processing end as a whole, and its weak signal detection capability and noise suppression capability in complex environments lay the foundation for high-precision identification of the source of the sound signal.

[0048] According to the embodiments of the present invention, the present invention uses physical means to achieve high-precision identification of the source of weak sound signals under low signal-to-noise ratio conditions, effectively making up for the shortcomings of traditional sound signal source identification technology being limited by the quality of detection data, and has broad application scenarios.

[0049] The optical resonance acoustic sensor 110 may include one of the following: an optical resonance acoustic sensor including a whispering gallery mode microcavity or a photonic crystal microcavity.

[0050] According to an embodiment of the present invention, an optical resonant acoustic sensor including a whispering gallery mode microcavity, a photonic crystal microcavity, etc. has the characteristics of a high quality factor and a small mode volume, which can greatly enhance the interaction between the detection light generated by the optical resonant acoustic sensor and the sound signal, thereby enabling the extraction of weak sound signals.

[0051] The modulation signal detection unit 120 can also be used to split and phase-shift the initial reference light to obtain two reference lights.

[0052] The modulated signal detection unit 120 may include a first beam splitter 121 , a second beam splitter 122 , a first phase shifter 123 , a second phase shifter 124 , a third beam splitter 125 , and a fourth beam splitter 126 .

[0053] The first beam splitter 121 can be used to split the acoustic modulated signal light into a first signal light and a second signal light.

[0054] According to an embodiment of the present invention, the first beam splitter 121 may be a 1×2 beam splitter.

[0055] According to an embodiment of the present invention, the first signal light It can be expressed as formula (3), the second signal light It can be expressed as formula (4).

[0056] (3).

[0057] (4).

[0058] The second beam splitter 122 can be used to split the initial reference light into a first path of initial sub-light and a second path of initial sub-light.

[0059] According to an embodiment of the present invention, the second beam splitter 122 may be a 1×2 beam splitter.

[0060] According to an embodiment of the present invention, the light field of the initial reference light It can be expressed as formula (5).

[0061] (5).

[0062] Among them, B j is the amplitude of the jth tooth of the initial reference light optical frequency comb, f r0 is the reference frequency of the initial reference light optical frequency comb, f rr is the repetition frequency of the initial reference light optical frequency comb.

[0063] According to an embodiment of the present invention, the reference frequency f of the initial reference light optical frequency comb is r0 and the reference frequency f of the signal light optical frequency comb s0 The repetition frequency f of the initial reference light optical frequency comb can be the same or different. rr and the repetition frequency f of the signal light optical frequency comb sr different.

[0064] The first phase shifter 123 can be used to perform a first phase shift on the first initial sub-light to obtain a first reference light.

[0065] The second phase shifter 124 can be used to perform a second phase shift on the second initial sub-light to obtain a second reference light, wherein the phase of the second reference light is orthogonal to the phase of the first reference light.

[0066] According to an embodiment of the present invention, the two reference beams having orthogonal phases may include a first reference beam and a second reference beam. The first reference beam has a first phase, and the second reference beam has a second phase. The first phase is orthogonal to the second phase. For example, the first phase may be 0, and the second phase may be π / 2.

[0067] The third beam splitter 125 can be used to perform beat frequency processing on the first signal light and the first reference light to obtain a first beat frequency signal.

[0068] According to an embodiment of the present invention, the third beam splitter 125 may be a 2×2 beam splitter.

[0069] The fourth beam splitter 126 can be used to perform beat frequency processing on the second signal light and the second reference light to obtain a second beat frequency signal.

[0070] According to an embodiment of the present invention, the fourth beam splitter 126 may be a 2×2 beam splitter.

[0071] The modulation signal detection unit 120 may further include a first balanced photodetector 127 and a second balanced photodetector 128 .

[0072] The first balanced photodetector 127 can be used to perform differential detection on the first beat frequency signal to obtain a first detection result.

[0073] According to an embodiment of the present invention, when the first phase of the first reference light is 0, the light field of the first reference light is It can be expressed as formula (6).

[0074] (6).

[0075] According to the embodiment of the present invention, it can be known from formulas (5) and (6) that the frequency of the first reference light beam corresponding to the optical frequency comb tooth number j is equal to the frequency of the initial reference light beam corresponding to the optical frequency comb tooth number j, both of which are f r0 +jf rr The second predetermined frequency interval may be equal to the repetition frequency of the initial reference light optical frequency comb.

[0076] According to an embodiment of the present invention, the first reference light and the first signal light are beat by each other in the first beam splitter 125, and a light field is obtained at the transmission end of the first beam splitter 125. , the light field is obtained at the reflection end of the first beam splitter 125 , after differential detection, the first detection result is obtained , as shown in formula (7).

[0077] (7);

[0078] in, is the frequency of the detection signal corresponding to the optical frequency comb tooth number j, , is the repetition frequency difference, , is the reference frequency difference, .

[0079] The second balanced photodetector 128 can be used to perform differential detection on the second beat frequency signal to obtain a second detection result.

[0080] According to an embodiment of the present invention, when the first phase of the second reference light is π / 2, the light field of the second reference light can be expressed as formula (8).

[0081] (8).

[0082] According to the embodiment of the present invention, it can be seen from formulas (5) and (8) that the frequency of the second reference light beam corresponding to the optical frequency comb tooth number j is equal to the frequency of the initial reference light beam corresponding to the optical frequency comb tooth number j, and both are The second predetermined frequency interval may be equal to the repetition frequency of the initial reference light optical frequency comb.

[0083] According to an embodiment of the present invention, it can be seen from formula (5), formula (6) and formula (8) that the repetition frequency of the first reference light and the repetition frequency of the second reference light are the same as the repetition frequency of the initial reference light optical frequency comb.

[0084] According to an embodiment of the present invention, the second reference light and the second signal light are beat in the fourth beam splitter 126, and a light field is obtained at the transmission end of the fourth beam splitter 126. , the light field is obtained at the reflection end of the fourth beam splitter 126 , after differential detection, the second detection result is obtained , as shown in formula (9).

[0085] (9).

[0086] The data processing unit 130 may generate a voiceprint QR code based on the first detection result and the second detection result of the preset time period, which may include: dividing the preset time period into multiple sub-time periods; for each of the multiple sub-time periods, performing a fast Fourier transform on the first detection result in each sub-time period to obtain a first transformation result, and performing a fast Fourier transform on the second detection result in each sub-time period to obtain a second transformation result; calculating the square sum of the first transformation result and the second transformation result to obtain a square sum result; and generating a voiceprint QR code based on the square sum results in the multiple sub-time periods.

[0087] According to an embodiment of the present invention, the first transformation result and the second transformation result They can be expressed as formula (10) and formula (11) respectively.

[0088] (10);

[0089] (11);

[0090] in, For sub-periods The average value of is the impulse function.

[0091] According to an embodiment of the present invention, the plurality of sub-periods includes M sub-periods, where M is an integer greater than 1.

[0092] For example, M can be 3, 5, 10, 20, 30, 50, 100 or 1000, etc.

[0093] The data processing unit 130 calculates the square sum of the first transformation result and the second transformation result, and obtaining the square sum result may include: in each sub-period, for each detection frequency among the multiple detection frequencies, calculating the square sum of the first transformation result and the second transformation result corresponding to each detection frequency, and obtaining the square sum result corresponding to each detection frequency, wherein the multiple detection frequencies correspond one-to-one to the multiple detection lights, and the multiple detection frequencies are each in a frequency band corresponding to their respective resonant modes.

[0094] According to the embodiment of the present invention, according to formulas (7) and (9), the frequency f of the detection signal corresponding to the optical frequency comb tooth number j is j For the detection frequency Therefore, the detection signal corresponding to the optical frequency comb tooth number j can be The corresponding detection signal. Then the transformation result of the detection signal corresponding to the optical frequency comb tooth number j can be The corresponding transformation results include a first transformation result and a second transformation result.

[0095] According to an embodiment of the present invention, when calculating the square sum of the first transformation result and the second transformation result, the detection frequency The corresponding sum of squares is .

[0096] The data processing unit 130 generates a voiceprint QR code based on the square sum results in multiple sub-time periods, which may include: determining the data included in the mth row of the voiceprint QR code based on the square sum results corresponding to multiple detection frequencies in the mth sub-time period, where m is an integer greater than 1 and less than or equal to M.

[0097] According to an embodiment of the present invention, by determining the data included in the mth row of the voiceprint QR code based on the square sum results corresponding to multiple detection frequencies in the mth sub-time period, the detection results of each time period are integrated to generate a voiceprint QR code, so that the horizontal direction of the mth voiceprint QR code can reflect the sound signal information of the sound signal reflected in the multiple resonance modes of the optical resonant acoustic sensor, and the vertical direction can reflect the waveform characteristics of the sound signal such as amplitude, frequency, pause, etc.

[0098] The data processing unit 130 determines the target voiceprint QR code that matches the voiceprint QR code from the multiple preset voiceprint QR codes included in the voiceprint QR code database, which may include: calculating the similarities between the voiceprint QR code and the multiple preset voiceprint QR codes respectively to obtain multiple similarities; and determining the target voiceprint QR code from the multiple preset voiceprint QR codes according to the maximum similarity among the multiple similarities.

[0099] Figure 2 FIG. 4 shows a structural diagram of a voiceprint recognition system according to another embodiment of the present invention.

[0100] like Figure 2 As shown, the voiceprint recognition system 100 may include an optical resonant acoustic sensor 110, a modulation signal detection unit 120, a data processing unit 130, a first optical frequency comb 140, a polarization controller 150, and a second optical frequency comb 160. Figure 2 The modulation signal detection unit 120 and the data processing unit 130 in Figure 1 The modulation signal detection unit 120 and the data processing unit 130 have similar structures and functions, and for the sake of simplicity, they are not described here in detail.

[0101] The first optical frequency comb 140 can be used to output a plurality of signal lights.

[0102] The polarization controller 150 can be used to perform polarization modulation on a plurality of signal lights to obtain a plurality of polarization-modulated signal lights.

[0103] According to an embodiment of the present invention, multiple polarization-modulated signal lights are obtained by performing polarization modulation on multiple signal lights, so that the multiple polarization-modulated signal lights can excite multiple resonance modes of the optical resonance acoustic sensor 110 .

[0104] Figure 2 The optical resonant acoustic sensor 110 and Figure 1 The difference between the optical resonant acoustic sensor 110 is that: Figure 2 The optical resonance acoustic sensor 110 can also be used to be excited into multiple resonance modes under the action of multiple polarization modulated signal lights, and generate multiple detection lights respectively in the multiple resonance modes.

[0105] The second optical frequency comb 160 can be used to output initial reference light.

[0106] According to the embodiments of the present invention, the voiceprint recognition system provided by the embodiments of the present invention has the following advantages: First, it has the ability to detect weak sound signals. The optical resonance acoustic sensor in the voiceprint recognition system can detect weak sound signals in real time. From the perspective of working bandwidth and sound pressure sensitivity level, it far exceeds the indicators of optical fiber acoustic sensors and piezoelectric sensors. Second, it has the ability to suppress noise in complex environments. Because the sensing mechanism of the optical resonance acoustic sensor involves multiple resonance modes, the noise intensity of noise inconsistent with the sound signal will be greatly reduced in some modes. Therefore, in multiple modes, a sound signal with reduced noise can be obtained. Third, signal identification capability. Existing solutions use the data end for processing, which is limited by the detection capability of data quality. The present invention combines the physical end and the data end as a whole. Its weak signal detection capability and noise suppression capability in complex environments lay the foundation for identifying the source of the sound signal.

[0107] Figure 3 A flow chart of a voiceprint recognition method according to an embodiment of the present invention is shown.

[0108] like Figure 3 As shown, the voiceprint recognition method may include operations S310 to S330.

[0109] According to an embodiment of the present invention, Figure 3 The voiceprint recognition method in is applied to the above-mentioned voiceprint recognition system.

[0110] In operation S310 , the optical resonance acoustic sensor detects an acoustic signal from a target scene using a plurality of detection lights in a plurality of resonance modes to obtain acoustic modulated signal light.

[0111] In operation S320, the modulation signal detection unit beats the two phase-orthogonal reference lights with the acoustic modulation signal light respectively to obtain two beat signals, and performs differential detection on the two beat signals to obtain a first detection result and a second detection result, wherein the detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval.

[0112] In operation S330, the data processing unit generates a voiceprint QR code based on the first detection result and the second detection result of the preset time period; determines a target voiceprint QR code that matches the voiceprint QR code from a plurality of preset voiceprint QR codes included in the voiceprint QR code database; and determines an object that matches the sound signal based on the target voiceprint QR code.

[0113] It should be noted that the voiceprint recognition method part in the embodiment of the present invention corresponds to the voiceprint recognition system part in the embodiment of the present invention. The description of the voiceprint recognition method part specifically refers to the voiceprint recognition system part, which will not be repeated here.

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0115] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present invention is defined by the accompanying embodiments and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present invention.

Claims

1. A voiceprint recognition system, characterized in that: include: An optical resonance acoustic sensor is used to detect sound signals from a target scene using multiple detection lights in multiple resonance modes to obtain acoustically modulated signal light; a modulation signal detection unit, configured to perform beat frequency detection on two reference lights having orthogonal phases with the acoustic modulation signal light to obtain two beat frequency signals, and to perform differential detection on the two beat frequency signals to obtain a first detection result and a second detection result, wherein the detection frequencies of the plurality of detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the plurality of reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval; a data processing unit, configured to divide a preset time period into a plurality of sub-time periods; and for each of the plurality of sub-time periods, perform a fast Fourier transform on the first detection result in each sub-time period to obtain a first transformation result, and perform a fast Fourier transform on the second detection result in each sub-time period to obtain a second transformation result; The plurality of sub-periods include M, where M is an integer greater than 1; in each sub-period, for each detection frequency among a plurality of detection frequencies, calculating a square sum of a first transformation result and a second transformation result corresponding to each detection frequency to obtain a square sum result corresponding to each detection frequency, wherein the plurality of detection frequencies correspond one-to-one to the plurality of detection lights; Determine the data included in the mth row of the voiceprint QR code according to the sum of squares corresponding to the multiple detection frequencies in the mth sub-period, where m is an integer greater than 1 and less than or equal to M; A target voiceprint QR code matching the voiceprint QR code is determined from a plurality of preset voiceprint QR codes included in a voiceprint QR code database; and an object matching the sound signal is determined based on the target voiceprint QR code.

2. The system according to claim 1, wherein: The system further comprises: a first optical frequency comb, configured to output a plurality of signal lights; a polarization controller, configured to perform polarization modulation on the plurality of signal lights to obtain a plurality of polarization-modulated signal lights; The optical resonance acoustic sensor is further configured to be excited into the multiple resonance modes under the action of the multiple polarization modulated signal lights, and to generate the multiple detection lights respectively in the multiple resonance modes.

3. The system according to claim 2, characterized in that The system further comprises: a second optical frequency comb for outputting an initial reference light; The modulation signal detection unit is further configured to perform light splitting and phase shifting on the initial reference light to obtain the two reference lights.

4. The system according to claim 3, characterized in that The modulation signal detection unit includes: a first beam splitter, configured to split the acoustically modulated signal light into a first signal light and a second signal light; a second beam splitter, configured to split the initial reference light into a first path of initial sub-light and a second path of initial sub-light; a first phase shifter, configured to perform a first phase shift on the first initial sub-light to obtain a first reference light; a second phase shifter, configured to perform a second phase shift on the second initial sub-light beam to obtain a second reference light beam, wherein the phase of the second reference light beam is orthogonal to the phase of the first reference light beam; a third beam splitter, configured to perform a beat frequency operation on the first signal light and the first reference light to obtain a first beat frequency signal; The fourth beam splitter is used to perform beat frequency analysis on the second signal light and the second reference light to obtain a second beat frequency signal.

5. The system according to claim 4, characterized in that The modulation signal detection unit further includes: a first balanced photodetector, configured to perform differential detection on the first beat frequency signal to obtain the first detection result; The second balanced photodetector is configured to perform differential detection on the second beat frequency signal to obtain the second detection result.

6. The system according to claim 4, characterized in that The optical resonance acoustic sensor includes one of the following: an optical resonance acoustic sensor including a whispering gallery mode microcavity or a photonic crystal microcavity.

7. The system according to claim 1, wherein: The data processing unit determines, from a plurality of preset voiceprint QR codes included in the voiceprint QR code database, a target voiceprint QR code that matches the voiceprint QR code, including: Calculating similarities between the voiceprint QR code and the plurality of preset voiceprint QR codes respectively to obtain a plurality of similarities; The target voiceprint QR code is determined from the multiple preset voiceprint QR codes according to the maximum similarity among the multiple similarities.

8. A voiceprint recognition method, applied to the system according to any one of claims 1 to 7, characterized in that: The method comprises: The optical resonance acoustic sensor uses multiple detection lights in multiple resonance modes to detect sound signals from the target scene and obtain acoustically modulated signal light; The modulation signal detection unit performs beat frequency detection on two reference lights with orthogonal phases with the acoustic modulation signal light to obtain two beat frequency signals, and performs differential detection on the obtained two beat frequency signals to obtain a first detection result and a second detection result, wherein the detection frequencies of the multiple detection lights are linearly arranged based on a first predetermined frequency interval, and the frequencies of the multiple reference sub-lights included in the reference light are linearly arranged based on a second predetermined frequency interval, and the second predetermined frequency interval is different from the first predetermined frequency interval; The data processing unit divides a preset time period into a plurality of sub-time periods; for each of the plurality of sub-time periods, performs a fast Fourier transform on the first detection result in each sub-time period to obtain a first transformation result, and performs a fast Fourier transform on the second detection result in each sub-time period to obtain a second transformation result; The plurality of sub-periods include M, where M is an integer greater than 1; in each sub-period, for each detection frequency among a plurality of detection frequencies, calculating a square sum of a first transformation result and a second transformation result corresponding to each detection frequency to obtain a square sum result corresponding to each detection frequency, wherein the plurality of detection frequencies correspond one-to-one to the plurality of detection lights; According to the square sum results corresponding to the multiple detection frequencies in the mth sub-time period, determine the data included in the mth row of the voiceprint QR code, where m is an integer greater than 1 and less than or equal to M; from the multiple preset voiceprint QR codes included in the voiceprint QR code database, determine the target voiceprint QR code that matches the voiceprint QR code; based on the target voiceprint QR code, determine the object that matches the sound signal.

Citation Information

Patent Citations

  • Identity recognition method and device

    CN113035202A

  • Interaction-unlimited speech enhancement method, system and terminal based on ultrasonic sensing

    CN117935825A