A method, system, electronic device, and readable storage medium for determining a sound source.
By acquiring real-time audio information in multi-person scenarios for audio recognition and activity detection, and combining multi-directional sound receiving components and infrared imaging, the problem of insufficient adaptability and stability of sound recognition technology in multi-person scenarios is solved, and fast and accurate sound source localization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing voice recognition technologies have low adaptability and stability to environmental changes in multi-person scenarios, resulting in poor versatility.
By acquiring real-time audio information, performing audio recognition and activity detection, utilizing multi-directional sound receiving components and sound source localization parameters, the location of the sound source can be determined, and combined with infrared imaging to obtain posture information, thus achieving rapid and accurate sound source localization.
In multi-person scenarios, it enables the rapid and accurate acquisition of the location information of a specific voice-speaking object, improving the versatility and stability of voice recognition technology.
Smart Images

Figure CN116504272B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound source localization, and more specifically, to a sound source determination method, system, electronic device, and readable storage medium. Background Technology
[0002] Voice recognition technology can identify human voices, snoring, abnormal noises, and the sounds of moving objects. Therefore, it can be widely used in speech processing, fault detection, and other fields. However, voice recognition technology is only applicable to a single object and has high requirements for the pronunciation quality of the object being recognized. This means that voice recognition technology can only adapt to situations with low degree of environmental change, thus reducing its versatility and stability. Summary of the Invention
[0003] The embodiments of this application provide a sound source determination method, system, electronic device, and readable storage medium, which can be applied to sound recognition in multi-person scenarios. This solves the problem that sound recognition technology in related technologies can only adapt to situations with low degree of environmental change, resulting in low versatility and stability of sound recognition technology.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, a method for determining a sound source is provided. The method includes: acquiring initial audio information acquired in real time; performing audio recognition processing on the initial audio information to obtain an audio recognition result; if the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as target audio information; performing audio information activity detection on the target audio information to obtain target audio activity information; and performing sound source localization on the sound-emitting object corresponding to the target audio activity information according to the sound source localization parameters corresponding to the target audio activity information to obtain target location information of the sound-emitting object.
[0006] Optionally, obtaining the initial audio information acquired in real time includes: setting up sound receiving components in at least two different locations, and acquiring audio information in real time based on the sound receiving components in at least two different locations to obtain the initial audio information.
[0007] Optionally, the preset audio recognition conditions include a preset audio signal frequency range and a preset audio signal sound pressure range. Performing audio recognition processing on the initial audio information to obtain an audio recognition result includes: performing frequency feature recognition or sound pressure feature recognition on the initial audio information to obtain an audio recognition result; if the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as target audio information includes: if the audio recognition result indicates that the frequency of the initial audio information meets the preset audio signal frequency range and the sound pressure of the initial audio information meets the preset audio signal sound pressure range, then using the initial audio information corresponding to the audio recognition result as target audio information.
[0008] Optionally, the step of detecting audio activity in the target audio information to obtain target audio activity information includes: acquiring the audio signal amplitude and zero-crossing rate corresponding to the target audio information, wherein the zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses zero, and the sampling information is obtained after sampling the target audio information multiple times; comparing the audio signal amplitude with a preset amplitude threshold to obtain an amplitude threshold comparison result; comparing the zero-crossing rate with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; and determining the target audio information as the target audio activity information when the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, and the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0009] Optionally, the step of detecting audio activity in the target audio information to obtain target audio activity information includes: acquiring the audio signal amplitude or zero-crossing rate corresponding to the target audio information, wherein the zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses zero, and the sampling information is obtained after sampling the target audio information multiple times; comparing the audio signal amplitude with a preset amplitude threshold to obtain an amplitude threshold comparison result; or comparing the zero-crossing rate with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; and determining the target audio information as the target audio activity information when the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, or the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0010] Optionally, the sound source localization parameters include the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information. Based on the sound source localization parameters corresponding to the target audio activity information, the sound source is localized for the sound-emitting object corresponding to the target audio activity information to obtain the target position information of the sound-emitting object. This includes: performing angular localization on the sound-emitting object based on the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information to obtain the azimuth angle parameters between the sound-emitting object and each of the receiving components; performing distance localization estimation on the sound-emitting object based on the audio signal propagation speed and the time difference between the audio signals received by each of the receiving components in the preceding and following periods to obtain the straight-line distance parameters between the sound-emitting object and the receiving components; and determining the relative position information between the sound-emitting object and the receiving components in three-dimensional space based on the azimuth angle parameters and the straight-line distance parameters, and using the relative position information as the target position information.
[0011] Optionally, the image acquisition area corresponding to the infrared image acquisition module is adjusted to obtain a target image acquisition area, wherein the target image acquisition area includes the sound-emitting object;
[0012] The infrared image acquisition module performs an infrared image capture operation on the target image acquisition area to obtain the infrared image posture information of the sound-emitting object, and stores the infrared image posture information, which is used to obtain the corresponding posture correction method.
[0013] According to one aspect of the embodiments of this application, a sound source determination system is provided. The system includes a sound receiving module for acquiring initial audio information collected in real time; a processing module for performing audio recognition processing on the initial audio information to obtain an audio recognition result; and, when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as target audio information; a detection module for performing audio information activity detection on the target audio information to obtain target audio activity information; and a positioning module for performing sound source positioning on the sound source emitting object corresponding to the target audio activity information based on the sound source positioning parameters corresponding to the target audio activity information to obtain target location information of the sound emitting object.
[0014] According to one aspect of the embodiments of this application, an electronic device is provided, including one or more processors; and a storage device for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to perform the method as described above.
[0015] According to one aspect of the embodiments of this application, an embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform the method described above.
[0016] According to one aspect of the embodiments of this application, an embodiment of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the methods described above in the various embodiments.
[0017] The technical solution provided in the embodiments of this application acquires initial audio information collected in real time; performs audio recognition processing on the initial audio information to obtain an audio recognition result; when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, the initial audio information corresponding to the audio recognition result is taken as target audio information; audio information activity detection is performed on the target audio information to obtain target audio activity information; and the sound source localization is performed on the sound source localization parameter corresponding to the target audio activity information to obtain the target position information of the sound source localization object. Wherein, when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, it indicates that the sound source localization object corresponding to the initial audio information is a specific sound source localization object. After obtaining the target audio activity information through audio information activity detection based on the target audio information, the sound source localization of the sound source localization object is directly performed based on the target audio information, thus achieving rapid and accurate acquisition of the target position information of a specific sound source localization object, avoiding the problem in related technologies where it is impossible to localize a specific sound source localization object.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0020] Figure 1 This is a basic flowchart illustrating an exemplary embodiment of the sound source determination method of this application;
[0021] Figure 2 This is a basic structural diagram of a sound source determination system illustrated in an exemplary embodiment of this application;
[0022] Figure 3 This is a basic structural diagram of another sound source determination system shown in an exemplary embodiment of this application;
[0023] Figure 4 This is a basic flowchart illustrating an exemplary embodiment of the sound source determination method of this application;
[0024] Figure 5 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0026] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the accompanying diagrams are merely illustrative and do not necessarily include all content and operations, nor do they necessarily have to be executed in the described order. For example, some operations may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0028] It should also be noted that "multiple" as mentioned in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0029] Example 1
[0030] To address the aforementioned technical problems, embodiments of this application provide a method for determining a sound source, such as... Figure 1 As shown, the method includes:
[0031] S101. Obtain the initial audio information acquired in real time;
[0032] S102. Perform audio recognition processing on the initial audio information to obtain an audio recognition result; if the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, use the initial audio information corresponding to the audio recognition result as the target audio activity information.
[0033] S103. Perform audio information activity detection on the target audio information to obtain target audio activity information;
[0034] S104. Based on the sound source localization parameters corresponding to the target audio activity information, perform sound source localization on the sound source emitting object corresponding to the target audio activity information to obtain the target location information of the sound emitting object.
[0035] It is understood that the initial audio information mentioned above is digital information converted from real-time acquired sound. This real-time acquired sound can be from a single speaker or from multiple speakers. If it is determined that the received sound is from a single speaker, then the sound can be directly converted into digital information to obtain the initial audio information. If it is determined that the received sound is from multiple speakers, then the acquired sound needs to be separated, and the initial audio information corresponding to each speaker needs to be obtained based on the separated sound. This example does not limit the method of speech role separation; relevant personnel can flexibly choose speech role separation methods to perform speech role separation.
[0036] Continuing from the previous example, after the sound receiving component receives the sound, it obtains the analog signal of the sound. Since analog signals are highly variable and easily affected by external environmental interference, and are not conducive to subsequent signal reconstruction and signal processing, it is necessary to convert the analog signal into the corresponding digital signal to obtain the corresponding initial audio information.
[0037] This embodiment does not limit the method of converting analog signals into digital signals. The following is an example of converting analog signals into digital signals to obtain initial audio information: the conversion of analog signals into digital signals includes three main steps: sampling, quantization, and encoding.
[0038] Sampling: Digitizing the analog signal along the time axis at a certain sampling rate. Specifically, first, we take multiple points sequentially along the time axis at fixed time intervals T (assuming T = 0.1s) (points on the wave corresponding to 1-10 in the figure). T is called the sampling period, and the reciprocal of T is the target frequency for this sampling (f = 1 / T = 10Hz), where f represents the number of samples per second, measured in Hertz (Hz). Clearly, the higher the target frequency and the more sampling points per unit time, the better the original waveform can be represented (if numerous points are collected at high frequency and density, it is equivalent to completely recording the original waveform). It's understandable that the higher the target frequency and the more sampling points, the better the original waveform can be represented. For a more detailed explanation, refer to the Nyquist sampling theorem: the target frequency f must be greater than twice the maximum vibration frequency fmax of the original audio signal (i.e., f > 2 * fmax, where fmax is called the Nyquist frequency) for the sampling results to be used to completely reconstruct the original audio signal; if the sampling rate is lower than 2 * fmax, then audio sampling will be distorted. For example, if you want to sample the original audio with the highest frequency fmax = 8kHz, then the target frequency f must be at least 16kHz.
[0039] It is understood that the specific sampling rate can be flexibly set by relevant personnel according to actual usage needs, and this embodiment does not limit it.
[0040] The second step is quantization: digitizing the analog signal with a certain precision along the amplitude axis. Specifically, after sampling, we proceed to the second step of audio digitization: quantization. Sampling digitizes the audio signal along the time axis, obtaining multiple sample points; while quantization digitizes the signal along the amplitude direction, obtaining the amplitude value of each sample point.
[0041] The third step is encoding: recording the sampled / quantized data in a specific format. Specifically, after quantization, we obtain the amplitude value of each sample point. Next comes the final step in digitizing the audio signal: encoding. Encoding converts the amplitude quantization value of each sample point into a binary byte sequence that a computer can understand.
[0042] Referring to the table in the encoding section, the sample number indicates the sampling order, and the sample value (decimal) is the quantized amplitude value. The sample value (binary) is the encoded data after amplitude conversion. Ultimately, we obtain a binary byte sequence in the form of "0" and "1", which is a discrete digital signal. What we obtain here is the uncompressed raw audio sample data stream, also called PCM audio data (Pulse Code Modulation). In practical applications, other encoding algorithms are often used for further compression; this embodiment does not limit this.
[0043] The collected sound can be the sound produced by the breathing of the voice-generating object. By collecting the breathing sound and converting it into a digital signal to obtain initial audio information, and when the target audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, the voice-generating object corresponding to the initial audio information is located according to the sound source localization parameters corresponding to the initial audio information to obtain the target location information of the voice-generating object. When the target audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, it indicates that the voice-generating object corresponding to the initial audio information is a specific voice-generating object. By directly locating the sound source of the voice-generating object based on the initial audio information, the location information of a specific voice-generating object can be obtained quickly, accurately and stably in a multi-person scenario, avoiding the problem in related technologies that it is impossible to locate a specific voice-generating object.
[0044] In some examples of this embodiment, acquiring the initial audio captured in real time includes:
[0045] A sound receiving component is set up in at least two different locations, and audio is collected in real time based on the sound receiving component in at least two different locations to obtain the initial audio information.
[0046] The sound-receiving component includes, but is not limited to, a microphone or a terminal device equipped with a microphone (e.g., a mobile phone, a smartwatch, etc.).
[0047] The following example uses a microphone as a sound receiving component. Microphones are set at at least two different locations in an area, and audio information is collected in real time through at least two microphones simultaneously. That is, when collecting audio information, at least two microphones work synchronously, so that the collected audio information is more complete. The two microphones mentioned above are the left channel microphone and the right channel microphone set on the PCB board.
[0048] It is understood that this embodiment does not limit the number or location of the microphones. The microphones can be placed in at least two different locations in a region. For example, there are two microphones, one on the left and the other on the right; or three microphones, two on the left and the other on the right; or four microphones, one on the east, one on the south, one on the west and one on the north.
[0049] In some examples of this embodiment, the preset audio recognition conditions include a preset audio signal frequency range and a preset audio signal sound pressure range. Performing audio recognition processing on the initial audio information to obtain a target audio recognition result includes: performing frequency feature recognition or sound pressure feature recognition on the initial audio information to obtain an audio recognition result; when the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as the target audio information includes: if the audio recognition result indicates that the frequency of the initial audio information meets the preset audio signal frequency range and the sound pressure of the initial audio information meets the preset audio signal sound pressure range, then using the initial audio information corresponding to the audio recognition result as the target audio information.
[0050] In the step of performing frequency feature recognition or sound pressure feature recognition on the initial audio information to obtain the initial audio recognition result, the audio signal frequency or audio signal sound pressure of the initial audio information is first obtained. Then, the frequency of the initial audio information is matched with a preset audio signal frequency range, and the sound pressure of the initial audio information is matched with a preset audio signal sound pressure range. If the frequency of the initial audio information is within the preset audio signal frequency range and the sound pressure of the initial audio information is within the preset audio signal sound pressure range, then it is determined that the initial audio information meets the preset audio recognition conditions indicated by the audio recognition result, and the initial audio information corresponding to the audio recognition result is taken as the target audio information. Conversely, if the frequency of the initial audio information is not within the preset audio signal frequency range or the sound pressure of the initial audio information is not within the preset audio signal sound pressure range, it is determined that the initial audio information does not meet the preset audio recognition conditions indicated by the audio recognition result, and the initial audio information corresponding to the audio recognition result is not taken as the target audio information.
[0051] In some examples, if the audio recognition result indicates that the frequency of the initial audio information meets the preset audio signal frequency range or the sound pressure of the initial audio information meets the preset audio signal sound pressure range, the initial audio information corresponding to the audio recognition result is used as the target audio information.
[0052] It is understandable that the preset audio signal frequency range and preset audio signal sound pressure range are determined by relevant personnel based on actual usage needs. For example, if relevant personnel want to identify a person snoring, they will obtain the snoring composite audio frequency and then determine the preset audio signal frequency range based on the snoring composite audio frequency; they will also obtain the sound pressure intensity of snoring sounds from different areas of the larynx and then determine the preset audio signal sound pressure range based on the sound pressure intensity of snoring sounds from different areas of the larynx.
[0053] Continuing from the previous example, the sound pressure levels of snoring sounds emitted from different areas of the larynx are as follows:
[0054] (1) The average sound pressure level of snoring in the soft palate was 36.58 (Decibel-Sound Pressure Level, dBSPL).
[0055] (2) The average sound pressure level of snoring from the epiglottis was 23.49 dBSPL;
[0056] (3) The average sound pressure intensity of snoring at the root of the tongue is 16.3 dBSPL.
[0057] To avoid errors, an error margin of 3 dBSPL is set, resulting in preset audio signal sound pressure ranges of 33.58 dBSPL to 39.58 dBSPL, 33.58 dBSPL to 39.58 dBSPL, and 13.3 dBSPL to 16.3 dBSPL. The initial audio information can be a complex snoring sound, including various snoring audio frequencies. If the initial audio information contains sound pressure levels falling within the range of 33.58 dBSPL to 39.58 dBSPL, and the initial audio information contains sound pressure levels falling within the range of 33.58 dBSPL to 39.58 dBSPL, and the initial audio information contains sound pressure levels falling within the range of 33.58 dBSPL to 39.58 dBSPL, and the initial audio information contains sound pressure levels falling within the range of 33.58 dBSPL to 39.58 dBSPL, then the initial audio information's sound pressure level falls within the range of 33.58 dBSPL to 39.58 dBSPL. If the initial audio information has a sound pressure level between 13.3 dBSPL and 16.3 dBSPL, then it is determined that the sound pressure level is within the preset audio signal sound pressure level range. Conversely, if the sound pressure level of the initial audio information does not meet any of the above three conditions, that is, if the sound pressure level of the initial audio information does not fall within the range of 33.58 dBSPL to 39.58 dBSPL, or the sound pressure level of the initial audio information does not fall within the range of 13.3 dBSPL to 16.3 dBSPL, then it is determined that the sound pressure level of the initial audio information is not within the preset audio signal sound pressure level range.
[0058] Similarly, if the frequency of the acquired snoring composite audio is 160-190Hz, then 160-190Hz is used as the preset audio signal frequency range. Then the frequency of the initial audio information is determined. If the frequency of the initial audio information falls within the 160-190Hz range, then the frequency of the initial audio information is determined to be within the preset audio signal frequency range; otherwise, the frequency of the initial audio information is determined to be outside the preset audio signal frequency range.
[0059] It is understood that relevant personnel may set multiple preset audio signal frequency ranges and multiple preset audio signal sound pressure ranges according to actual usage needs, and this embodiment does not limit this.
[0060] In some examples of this embodiment, the step of performing audio information activity detection on the target audio information to obtain target audio activity information includes:
[0061] Obtain the audio signal amplitude and zero-crossing rate corresponding to the initial audio information;
[0062] The amplitude of the audio signal is compared with a preset amplitude threshold to obtain the amplitude threshold comparison result;
[0063] The zero-crossing rate is compared with the preset zero-crossing rate threshold to obtain the zero-crossing rate comparison result;
[0064] If the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, and the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold, the target audio information is determined as the target audio activity information.
[0065] The audio signal amplitude represents the volume of the sound corresponding to the initial audio information. The larger the audio signal amplitude of the target audio information, the louder the sound; conversely, the smaller the audio signal amplitude of the target audio information, the quieter the sound. The preset amplitude threshold can be set by relevant personnel according to actual usage needs. The audio signal amplitude is compared with the preset amplitude threshold to obtain the amplitude threshold comparison result. If the audio signal amplitude is higher than the preset amplitude threshold, it is determined that the audio signal amplitude of the target audio information meets the amplitude threshold, and the target audio information is identified as the target audio activity information. Conversely, if the audio signal amplitude is not higher than the preset amplitude threshold, it is determined that the audio signal amplitude of the target audio information does not meet the amplitude threshold, and the target audio information will not be identified as the target audio activity information.
[0066] Zero Crossing Rate (ZCR) is the number of times a sampled audio signal crosses a zero point in each frame of the audio signal. A preset zero crossing rate threshold can be determined based on the number of periodic signals from the sampled data. Specifically, the periodic signals from the sampled data can be directly used as the preset zero crossing rate threshold. It is understood that this example does not limit the determination of the preset zero crossing rate threshold to solely based on the number of periodic signals from the sampled data; relevant personnel can flexibly choose the method for determining the preset zero crossing rate threshold. By comparing the zero crossing rate with the preset zero crossing rate threshold, a zero crossing rate comparison result is obtained. If the zero crossing rate is lower than the preset zero crossing rate threshold, it indicates that the target audio information is usable normal audio, and in this case, the target audio information is identified as the target audio activity information. If the zero crossing rate is higher than the zero crossing rate threshold, it indicates that the target audio information is noise, and in this case, the target audio information will not be identified as the target audio activity information.
[0067] Continuing with the previous example, taking the periodic signal of the sampled number as a preset zero-crossing rate threshold as an example, after obtaining the zero-crossing rate of the target audio information, the zero-crossing rate of the target audio information is compared with the periodic signal of the sampled number of the target audio information; if the number of periodic signals of the sampled number is greater than the zero-crossing rate, then the target audio information is determined to be the target audio activity information; conversely, if the number of periodic signals is less than the zero-crossing rate, then the target audio information is determined to be noise, and the target audio information will not be determined to be the target audio activity information.
[0068] In some examples, the step of detecting audio activity in the target audio information to obtain target audio activity information includes: acquiring the audio signal amplitude or zero-crossing rate corresponding to the target audio information, wherein the zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses zero, and the sampling information is obtained after sampling the target audio information multiple times; comparing the audio signal amplitude with a preset amplitude threshold to obtain an amplitude threshold comparison result; or comparing the zero-crossing rate with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; and determining the target audio information as the target audio activity information when the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, or the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0069] It is understandable that, in some examples, only the audio signal amplitude corresponding to the initial audio information can be obtained. If the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, the target audio information is identified as the target audio activity information; otherwise, the target audio information will not be identified as the target audio activity information. In some examples, only the zero-crossing rate corresponding to the initial audio information can be obtained. If the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold, the target audio information is identified as the target audio activity information; otherwise, the target audio information will not be identified as the target audio activity information. In some examples, it is necessary to identify the target audio information as the target audio activity information if the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, and the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0070] In some examples of this embodiment, the sound source localization parameters include the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information. Based on the sound source localization parameters corresponding to the target audio activity information, the sound source is localized to the sound-emitting object corresponding to the target audio activity information to obtain the target position information of the sound-emitting object. This includes: performing angular localization on the sound-emitting object based on the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information to obtain the azimuth angle parameters between the sound-emitting object and each of the receiving components; performing distance localization estimation on the sound-emitting object based on the audio signal propagation speed and the time difference between the audio signals received by each of the receiving components in the preceding and following periods to obtain the straight-line distance parameters between the sound-emitting object and the receiving components; and determining the relative position information between the sound-emitting object and the receiving components in three-dimensional space based on the azimuth angle parameters and the straight-line distance parameters, and using the relative position information as the target position information.
[0071] Continuing the previous example, taking the case where there are sound receiving components on the left and right sides respectively, the sound receiving components on the left and right sides are separated by a distance, resulting in different timing of the received audio signals. There is a time difference in the received signal between the leftmost sound receiving component and the rightmost sound receiving component. Based on the time difference of the received audio signal between the two sound receiving components, the interval distance, the speed of sound propagation, and the sampling rate, the azimuth parameter of the sound-emitting object corresponding to the audio signal can be estimated. Based on the time difference of the received audio signal in the preceding and following periods and the speed of sound propagation, the straight-line distance between the sound-emitting object corresponding to the audio signal and each sound receiving component can be estimated. After obtaining the straight-line distance parameter and the azimuth parameter, the relative position information between the sound-emitting object and the sound receiving component in three-dimensional space is determined based on the straight-line distance parameter and the azimuth parameter, and the relative position information is used as the target position information.
[0072] In some examples of this embodiment, after locating the sound source of the object corresponding to the target audio activity information and obtaining the target location information of the sound-emitting object, the method further includes: adjusting the image acquisition area corresponding to the infrared image acquisition module according to the target location information to obtain a target image acquisition area, wherein the target image acquisition area includes the sound-emitting object; performing an infrared image shooting operation on the target image acquisition area through the infrared image acquisition module to obtain the infrared image posture information of the sound-emitting object, and storing the infrared image posture information, which is used to obtain a corresponding posture correction method. That is, according to the target location information of the sound-emitting object, the image acquisition area of the infrared image acquisition module is adjusted to include the sound-emitting object, so that the infrared image acquisition module can capture the sound-emitting object and obtain the infrared image posture information of the sound-emitting object. It is understandable that the infrared image acquisition module can perform infrared image capture operations on the target image acquisition area under sufficient light conditions to obtain the infrared image posture information of the sound-emitting object; at the same time, due to the characteristics of infrared light, the infrared image acquisition module can also perform infrared image capture operations on the target image acquisition area under dark (insufficient light) conditions to obtain the infrared image posture information of the sound-emitting object, making this solution applicable to situations with high degree of light variation.
[0073] Taking the determination of preset audio signal frequency range and preset audio signal sound pressure range based on the composite audio frequency and sound pressure intensity of the snoring sound as an example, if the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, it indicates that the snoring object is currently snoring. After locating the sound source of the snoring object and obtaining the target location information of the snoring object, the camera image acquisition corresponding to the infrared image acquisition module is adjusted to obtain the target image acquisition area. Through the infrared image acquisition module, an infrared image shooting operation is performed on the target image acquisition area to obtain the infrared image posture information of the snoring object, and the infrared image posture information is stored.
[0074] In some examples of this embodiment, the infrared image posture information includes posture information to be corrected, which is the posture information corresponding to the sound-emitting object when it emits the target audio information. After storing the infrared image posture information, the method further includes: finding a posture correction method corresponding to the posture information to be corrected; and displaying posture reminder information corresponding to the posture correction method to the sound-emitting object.
[0075] Continuing with the previous example, taking a snoring person as an example, different postures will cause the snoring degree of the person to snore to be different. Therefore, based on the obtained posture information to be corrected, the corresponding posture correction method can be found, and posture reminder information corresponding to the posture correction method can be displayed to the person, so that the person can correct their posture according to the posture correction method.
[0076] The sound source determination method provided in this embodiment includes: acquiring initial audio information collected in real time; performing audio recognition processing on the initial audio information to obtain an audio recognition result; when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as target audio information; performing audio information activity detection on the target audio information to obtain target audio activity information; and performing sound source localization on the sound source object corresponding to the target audio activity information according to the sound source localization parameters corresponding to the target audio activity information to obtain the target location information of the sound source object. Wherein, when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, it indicates that the sound source object corresponding to the initial audio information is a specific sound source object. After performing audio information activity detection based on the target audio information to obtain the target audio activity information, the sound source is directly localized on the sound source object based on the target audio information, achieving rapid and accurate acquisition of the target location information of a specific sound source object, avoiding the problem in related technologies where it is impossible to locate a specific sound source object.
[0077] Example 2
[0078] Based on the same technical concept, this embodiment also provides a sound source determination system, such as... Figure 2 As shown, the system includes:
[0079] The audio receiving module 1 is used to acquire the initial audio information collected in real time;
[0080] Processing module 2 is used to perform audio recognition processing on the initial audio information to obtain an audio recognition result; when the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, the initial audio information corresponding to the audio recognition result is used as the target audio information.
[0081] Detection module 3 is used to perform audio information activity detection on the target audio information to obtain target audio activity information;
[0082] The positioning module 4 is used to locate the sound source of the sound source corresponding to the target audio activity information based on the sound source positioning parameters corresponding to the target audio activity information, so as to obtain the target location information of the sound source.
[0083] The sound source determination system also includes an infrared image acquisition module 5. The infrared image acquisition module 5 is used to adjust the image acquisition area corresponding to the infrared image acquisition module 5 according to the target position information to obtain a target image acquisition area, which includes the sound-emitting object. The infrared image acquisition module 5 performs an infrared image shooting operation on the target image acquisition area to obtain the infrared image posture information of the sound-emitting object. The infrared image posture information is used to obtain a corresponding posture correction method.
[0084] The process of acquiring initial audio information in real time includes: setting up sound receiving components in at least two different locations, and acquiring audio information in real time based on the sound receiving components in at least two different locations to obtain the initial audio information.
[0085] The preset audio recognition conditions include a preset audio signal frequency range and a preset audio signal sound pressure range. The initial audio information is processed to obtain an audio recognition result, including: performing frequency feature recognition or sound pressure feature recognition on the initial audio information to obtain an audio recognition result; if the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, the initial audio information corresponding to the audio recognition result is used as the target audio information, including: if the audio recognition result indicates that the frequency of the initial audio information meets the preset audio signal frequency range and the sound pressure of the initial audio information meets the preset audio signal sound pressure range, the initial audio information corresponding to the audio recognition result is used as the target audio information.
[0086] The step of detecting audio activity in the target audio information to obtain target audio activity information includes: acquiring the audio signal amplitude and zero-crossing rate corresponding to the target audio information, wherein the zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses the zero point, and the sampling information is obtained after sampling the target audio information multiple times; comparing the audio signal amplitude with a preset amplitude threshold to obtain an amplitude threshold comparison result; comparing the zero-crossing rate with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; and determining the target audio information as the target audio activity information when the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, and the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0087] In some examples, the step of detecting audio activity in the target audio information to obtain target audio activity information includes: acquiring the audio signal amplitude or zero-crossing rate corresponding to the target audio information, wherein the zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses zero, and the sampling information is obtained after sampling the target audio information multiple times; comparing the audio signal amplitude with a preset amplitude threshold to obtain an amplitude threshold comparison result; or comparing the zero-crossing rate with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; and determining the target audio information as the target audio activity information when the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, or the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold.
[0088] The sound source localization parameters include the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information. Based on the sound source localization parameters corresponding to the target audio activity information, the sound source of the emitting object corresponding to the target audio activity information is localized to obtain the target position information of the emitting object. This includes: performing angular localization on the emitting object based on the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information to obtain the azimuth angle parameters between the emitting object and each of the receiving components; performing distance localization estimation on the emitting object based on the audio signal propagation speed and the time difference between the audio signals received by each receiving component in the preceding and following periods to obtain the straight-line distance parameters between the emitting object and the receiving components; and determining the relative position information between the emitting object and the receiving components in three-dimensional space based on the azimuth angle parameters and the straight-line distance parameters, and using the relative position information as the target position information.
[0089] The method further includes, after locating the sound source of the object corresponding to the target audio activity information to obtain the target location information of the object, adjusting the image acquisition area corresponding to the infrared image acquisition module according to the target location information to obtain a target image acquisition area, wherein the target image acquisition area includes the object; and performing an infrared image shooting operation on the target image acquisition area through the infrared image acquisition module to obtain the infrared image posture information of the object, wherein the infrared image posture information is used to obtain a corresponding posture correction method.
[0090] It is understood that the sound source determination system specifically includes: a processor, an audio processor, a filter, an audio amplifier, an image sensor, a lens module, an infrared light-emitting diode module, an ambient light sensor, an infrared cut, a memory, two precision condenser microphones, etc. The above devices together constitute the sound receiving module 1, processing module 2, detection module 3, positioning module 4, and infrared image acquisition module 5.
[0091] It should be understood that the combination of various modules in the sound source determination system provided in this embodiment can realize the various steps of the sound source determination method described above, and achieve the same technical effect as the various steps of the sound source determination method, which will not be repeated here.
[0092] Example 3
[0093] To better understand the present invention, a more specific example is provided for illustration, wherein this example provides a sound source determination method, which is applied to a sound source determination system, such as... Figure 3 The sound source determination system shown includes: a processor, an audio processor, a filter, an audio amplifier, an image sensor, a lens module, an infrared light-emitting diode module, an ambient light sensor, an infrared cut, a memory (shown in the figure), and two precision condenser microphones. These components together form the sound source determination system's sound pickup module, processing module, detection module, positioning module, and imaging module.
[0094] The LENS module is connected to the image sensor, the image sensor is connected to the filter, the filter is connected to the processor, and the processor is connected to the memory. The processor is also connected to the ALS sensor, IRCUT, and IRLED module.
[0095] Specifically, an ambient light sensor (ALS) is a sensor that can detect the intensity of ambient light. An infrared filter (IR-CUT) blocks infrared light from entering, allowing only visible light to enter the camera or camcorder. This ensures image clarity and avoids image distortion or color difference caused by infrared interference. During the day, the IR filter allows visible light to enter the camera or camcorder, resulting in vividly colored images; at night or in low-light environments, the IR filter automatically switches to infrared-transmitting mode to capture black and white images, thus achieving optimal image quality. An infrared diode unit (IR-LED) is an electronic component that emits infrared light waves, mainly used in infrared communication, infrared remote controls, and infrared sensors. It converts electrical energy into infrared energy, and the emitted infrared light can propagate through the air and be received by a photosensitive element at the receiving end, thereby realizing functions such as communication, remote control, and detection. IR-LEDs have the characteristics of long emission distance, strong anti-interference ability, low power consumption, and small size, and are widely used in smart homes, industrial automation, security monitoring, and other fields. Specifically, the infrared diode unit includes an infrared emitting diode (IR Transmitter LED) and an infrared receiving diode (IR Receiver LED). The infrared emitting diode is used to emit infrared light, and the infrared receiving diode is used to identify infrared information and transmit the identified infrared information to the image sensor.
[0096] A lens module (LEN) is a type of lens module used in optical imaging, optical measurement, and optical communication. It typically consists of a set of optical elements such as lenses, filters, and apertures. By adjusting the relative positions and angles of these elements, optical parameters such as the direction of light propagation, focusing effect, and light intensity distribution can be changed, thereby enabling the processing and control of optical signals. LEN modules are widely used in various optical systems due to their simple structure, ease of adjustment, and reusability.
[0097] Specifically, the system includes an image sensor and filters for acquiring light signals from the area where the monitored object is located. It also includes an IR LED module, an ambient light sensor (ALS), and an infrared filter (IR-CUT) to assist the image sensor in acquiring light signals in low-light conditions. The IR LED module, ambient light sensor, and infrared filter sense the current ambient light conditions and send this information to the processor. The processor uses this information to control the operation of the IR LED module, ambient light sensor, and infrared filter, thereby adjusting the image sensor to function properly in low-light environments and acquire light signals.
[0098] The system consists of two microphones connected to their respective amplifiers, which in turn are connected to an A / D converter. The A / D converter is connected to an audio processor, and the audio processor is connected to another processor. Together, the two microphones, their corresponding amplifiers, the A / D converter, and the audio processor work to collect audio information.
[0099] Taking the sound source localization of snoring using the above-mentioned sound source determination method as an example, such as... Figure 4 As shown, Figure 4The diagram shows the flow chart of the sound source determination method provided in this example. First, the left and right high-precision capacitor microphones of this sound source determination system are respectively set on the left and right sides of the sound receiving area to receive the sound pressure level (dBSPL; Decibel-Sound Pressure Level) of the snoring from each area of the throat during snoring within the area to obtain audio information. The A / D converter then converts the audio information into digital signals according to the analog electrical signal. The sound pressure level of the snoring from each area of the throat during snoring is as follows: (1) The composite frequency of snoring is 160~190Hz; (2) The average sound pressure level of snoring from the soft palate is 36.58dBSPL; (3) The average sound pressure level of snoring from the epiglottis is 23.49dBSPL; (4) The average sound pressure level of snoring from the root of the tongue is 16.3dBSPL. After confirming that the sound pressure level is indeed snoring, the process of locating the source of the snoring begins. First, the received audio signal is sampled and processed from analog to digital. The amplitude of the audio signal is then checked to see if it exceeds a threshold value (Threshold Value Detection; TVD). The threshold value can be determined through multiple adjustments via training and can be dynamically adjusted based on massive data statistics. Next is the Zero Crossing Rate (ZCR) determination. ZCR calculation estimates the number of times the sampled signal crosses zero. If ZCR <= the number of cycles of the sampled signal, the sampled data is not noise. If ZCR > the number of cycles of the sampled signal, the sampled data is noise, because the left and right channel values can cause significant waveform differences, resulting in echo noise. When the captured sound source signal is less than the TVD, the snoring signal has ended. The two high-precision condenser microphones in the left and right channels are separated by a distance, resulting in different timing of the received snoring sound signals. The time difference between the sound from the leftmost microphone and the rightmost microphone is the largest. Based on the time difference of the audio signals received by the two microphones, the distance between the two microphones, the speed of sound, and the sampling rate, the azimuth of the snoring sound can be estimated. Based on the time difference and speed of sound of the received snoring audio, the straight-line distance between the snoring sound and the system can be estimated. Therefore, the azimuth and relative straight-line distance between the snoring sound and the snoring system can be estimated.
[0100] Once the location of the snoring sound source is determined, infrared thermal imaging is activated to capture and store the infrared thermal image of the snorer's sleeping posture. This eliminates the need to search for the snorer among multiple sleepers in a dark area. The captured infrared thermal image of the snorer's sleeping posture can serve as a reference for correction and for identifying non-invasive snoring treatments.
[0101] Example 4
[0102] Embodiments of this application also provide an electronic device, including one or more processors and a storage device, wherein the storage device is used to store one or more computer programs, which, when executed by one or more processors, cause the electronic device to implement the above-described sound source determination method.
[0103] Figure 5 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0104] It should be noted that, Figure 5 The computer system 1800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0105] like Figure 5 As shown, the computer system 1800 includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in read-only memory (ROM) 1802 or a program loaded from storage portion 1808 into random access memory (RAM) 1803. The RAM 1803 also stores various programs and data required for system operation. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.
[0106] In some embodiments, the following components are connected to the I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the I / O interface 1805 as needed. A removable medium 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1810 as needed so that computer programs read from it can be installed into the storage section 1808 as needed.
[0107] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by processor (CPU) 1801, it performs various functions defined in the system of this application.
[0108] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory, flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and a computer program.
[0110] The units or modules described in the embodiments of this application can be implemented in software or hardware, and can also be located in a processor. The names of these units or modules do not necessarily limit the specific unit or module itself.
[0111] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned sound source determination method. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0112] Another aspect of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the sound source determination method as described above in the various embodiments.
[0113] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0114] Other embodiments of this application will readily conceive of by considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0115] The above content is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.
Claims
1. A method for determining a sound source, characterized in that, The method for determining the sound source includes: Acquire initial audio information captured in real time; The initial audio information is subjected to audio recognition processing to obtain an audio recognition result, including: frequency feature recognition or sound pressure feature recognition of the initial audio information to obtain an audio recognition result; When the audio recognition result indicates that the initial audio information meets the preset audio recognition conditions, the initial audio information corresponding to the audio recognition result is taken as the target audio information, wherein the preset audio recognition conditions include a preset audio signal frequency range and a preset audio signal sound pressure range; Performing audio information activity detection on the target audio information to obtain target audio activity information includes: determining whether the target audio information is the target audio activity information based on the audio signal amplitude and / or zero-crossing rate of the target audio information; Based on the sound source localization parameters corresponding to the target audio activity information, the sound source localization is performed on the sound source emitting object corresponding to the target audio activity information to obtain the target location information of the sound emitting object; The method further includes, after locating the sound source of the sound-emitting object corresponding to the target audio activity information to obtain the target location information of the sound-emitting object: Based on the target location information, the image acquisition area corresponding to the infrared image acquisition module is adjusted to obtain the target image acquisition area, which includes the sound-emitting object. The infrared image acquisition module performs an infrared image capture operation on the target image acquisition area to obtain the infrared image posture information of the sound-emitting object, and stores the infrared image posture information. The infrared image posture information is used to obtain the corresponding posture correction method. The speaker is shown posture reminder information corresponding to the posture correction method, so that the speaker can correct their posture according to the posture correction method.
2. The method according to claim 1, characterized in that, Acquire the initial audio information captured in real time, including: A sound receiving component is set up in at least two different locations, and audio information is collected in real time based on the sound receiving component in at least two different locations to obtain the initial audio information.
3. The method according to claim 1, characterized in that, When the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, using the initial audio information corresponding to the audio recognition result as the target audio information includes: If the audio recognition result indicates that the frequency of the initial audio information meets the preset audio signal frequency range and the sound pressure of the initial audio information meets the preset audio signal sound pressure range, then the initial audio information corresponding to the audio recognition result is taken as the target audio information.
4. The method according to claim 1, characterized in that, The step of performing audio information activity detection on the target audio information to obtain target audio activity information includes: The amplitude and zero-crossing rate of the audio signal corresponding to the target audio information are obtained. The zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses the zero point. The sampling information is obtained by sampling the target audio information multiple times. The amplitude of the audio signal is compared with a preset amplitude threshold to obtain the amplitude threshold comparison result; The zero-crossing rate is compared with the preset zero-crossing rate threshold to obtain the zero-crossing rate comparison result; If the amplitude threshold comparison result indicates that the audio signal amplitude is greater than the amplitude threshold, and the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold, the target audio information is determined as the target audio activity information.
5. The method according to claim 1, characterized in that, The step of performing audio information activity detection on the target audio information to obtain target audio activity information includes: The amplitude or zero-crossing rate of the audio signal corresponding to the target audio information is obtained. The zero-crossing rate is the number of times the sampling information corresponding to the target audio information crosses the zero point. The sampling information is obtained by sampling the target audio information multiple times. The amplitude of the audio signal is compared with a preset amplitude threshold to obtain the amplitude threshold comparison result; Alternatively, the zero-crossing rate can be compared with the preset zero-crossing rate threshold to obtain a zero-crossing rate comparison result; If the amplitude threshold comparison result indicates that the amplitude of the audio signal is greater than the amplitude threshold, or if the zero-crossing rate comparison result indicates that the zero-crossing rate is less than or equal to the zero-crossing rate threshold, the target audio information is determined as the target audio activity information.
6. The method according to claim 2, characterized in that, The sound source localization parameters include the time difference between the audio signals received by each of the receiving components, the interval distance between the receiving components, the audio signal propagation speed, and the sampling rate of the target audio information. Based on the sound source localization parameters corresponding to the target audio activity information, the sound source localization is performed on the sound-emitting object corresponding to the target audio activity information to obtain the target location information of the sound-emitting object, including: Based on the time difference of the audio signals received by each of the audio receiving components, the interval between each of the audio receiving components, the audio signal propagation speed, and the sampling rate of the target audio information, the sound-emitting object is located at an angle to obtain the azimuth angle parameters between the sound-emitting object and each of the audio receiving components. Based on the propagation speed of the audio signal and the time difference between the audio signals received by each of the receiving components in the preceding and following periods, a distance positioning estimation is performed on the sound-emitting object to obtain the straight-line distance parameter between the sound-emitting object and the receiving component. Based on the azimuth parameter and the straight-line distance parameter, the relative position information between the sound-emitting object and the sound-receiving component in three-dimensional space is determined, and the relative position information is used as the target position information.
7. A sound source determination system, characterized in that, The system includes: The audio receiving module is used to acquire initial audio information in real time. The processing module is used to perform audio recognition processing on the initial audio information to obtain an audio recognition result: performing frequency feature recognition or sound pressure feature recognition on the initial audio information to obtain an audio recognition result; when the audio recognition result indicates that the initial audio information meets preset audio recognition conditions, the initial audio information corresponding to the audio recognition result is used as the target audio information, wherein the preset audio recognition conditions include a preset audio signal frequency range and a preset audio signal sound pressure range; The detection module is used to perform audio information activity detection on the target audio information to obtain target audio activity information: determining whether the target audio information is the target audio activity information based on the audio signal amplitude and / or zero-crossing rate of the target audio information; The positioning module is used to locate the sound source of the sound source corresponding to the target audio activity information based on the sound source positioning parameters corresponding to the target audio activity information, so as to obtain the target location information of the sound source. The sound source determination system further includes a reminder module, used for: After locating the sound source of the object corresponding to the target audio activity information to obtain the target location information of the object, the image acquisition area corresponding to the infrared image acquisition module is adjusted according to the target location information to obtain a target image acquisition area, which includes the object. The infrared image acquisition module then performs an infrared image capture operation on the target image acquisition area to obtain the infrared image posture information of the object, and stores this information. This infrared image posture information is used to obtain a corresponding posture correction method. Posture reminder information corresponding to the posture correction method is displayed to the object, enabling it to correct its posture according to the method.
8. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the processor of the electronic device, causes the electronic device to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Noise suppression based on correlation of sound in a microphone array
CN104412616A
Method and apparatus for directional enhancement of speech elements in noisy environments
US20070053522A1