Method and system for discriminating a target person in an environment
By setting up multiple microphone modules in the environment to monitor the reverberation coefficient, calculate the average value and determine the difference, the problem of not being able to determine the direction of the target when the target person is not speaking is solved, and the function of silently determining the direction of the target is realized.
Patent Information
- Application Number
- CN202310033646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-01-10
AI Technical Summary
In existing technologies, if the target person does not speak, it is impossible to determine whether the target person is present in the target direction in the environment.
By setting up multiple microphone modules in multiple directions of the environment, the reverberation coefficient in each direction is monitored, the average reverberation coefficient is calculated, and the difference between the reverberation coefficient in the target direction and the average value is used to determine whether a target person is present in the target direction.
Even if the target person does not speak, it can accurately determine whether the target person has appeared in the target direction, making up for the shortcomings of existing technology.
Smart Images

Figure CN116072148B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing, and more particularly to a method, system, computer device, and computer-readable storage medium for determining the direction of a person in an environment. Background Technology
[0002] In existing technologies, methods for identifying a person in a target direction within an environment require the target person to speak, and then the presence of the target person in that direction is determined based on the sound of their speech. If the target person does not speak, it is impossible to determine whether they are present in that direction. Summary of the Invention
[0003] The purpose of this application is to provide a method, system, computer device, and computer-readable storage medium for identifying a person in a target direction in an environment, in order to solve the following technical problem: if the target person does not speak, it is impossible to determine whether the target person is present in the target direction.
[0004] One aspect of this application provides a method for identifying a person in a target direction in an environment, wherein multiple microphone modules are provided in multiple directions of the environment, and the method for identifying a person in a target direction in the environment includes:
[0005] Multiple reverberation coefficients in multiple directions are monitored using the multiple microphone modules;
[0006] Based on the multiple reverberation coefficients in the multiple directions, determine the average reverberation coefficient in each direction within a preset time period;
[0007] The difference between the target reverberation coefficient in the target direction and the average value of the reverberation coefficient is used to determine whether a target person is present in the target direction.
[0008] Optionally, each microphone module includes a speaker and a microphone; the monitoring of multiple reverberation coefficients in the multiple directions via the multiple microphone modules includes:
[0009] The first ultrasonic signal is played through the speaker of each microphone module;
[0010] The second ultrasonic signal is acquired by the microphone of each microphone module. The second ultrasonic signal is the ultrasonic signal formed by the reflection of the first ultrasonic signal by the environment.
[0011] The speaker-to-microphone impulse response function of each microphone module is determined based on the first ultrasonic segment signal and the second ultrasonic segment signal.
[0012] Based on the impact response function, the reverberation coefficient in each direction is determined.
[0013] Optionally, determining the average reverberation coefficient for each direction within a preset time period based on the multiple reverberation coefficients for the multiple directions includes:
[0014] The reverberation coefficient in each direction is tracked and analyzed within the preset time period, and the average reverberation coefficient in each direction is determined within the preset time period.
[0015] Optionally, determining whether a target person is present in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficients includes:
[0016] If the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is greater than a threshold, it is determined that the target person appears in the target direction.
[0017] If the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is not greater than a threshold, it is determined that the target person does not appear in the target direction.
[0018] Optionally, the method for determining the direction of a person in an environment further includes:
[0019] The microphones of the multiple microphone modules are combined to form a microphone array, thereby creating multiple beams;
[0020] The system picks up sound signals within multiple beam ranges, the multiple beam ranges corresponding to multiple directions in the environment;
[0021] The target direction of the target person is determined based on the beam range of the sound signal.
[0022] Optionally, the method for determining the direction of a person in an environment further includes:
[0023] The difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is used as the first decision factor;
[0024] The magnitude of the sound signal within the multiple beam ranges is used as the second decision factor;
[0025] Based on the first decision factor and the second decision factor, a comprehensive decision is made to determine the target person's target direction.
[0026] Optionally, the method for determining the direction of a person in the environment further includes: guiding the array selection of speakers and the shape of microphone beams of the plurality of microphone modules according to the target direction of the target person.
[0027] One aspect of this application provides a system for identifying a person in a target direction in an environment, wherein multiple microphone modules are provided in multiple directions of the environment, and the system for identifying a person in a target direction in the environment includes:
[0028] A monitoring module is used to monitor multiple reverberation coefficients in multiple directions through the multiple microphone modules;
[0029] The determining module is used to determine the average value of the reverberation coefficients in each direction within a preset time period based on the multiple reverberation coefficients in the multiple directions.
[0030] The discrimination module is used to determine whether a target person is present in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient.
[0031] One aspect of this application provides a computer device, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for determining a person in a target direction in an environment as described above.
[0032] One aspect of this application provides a computer-readable storage medium storing a computer program that can be executed by at least one processor to perform the steps of the method for determining the direction of a person in an environment as described above.
[0033] The method, system, computer device, and computer-readable storage medium for determining the direction of a person in an environment provided in this application have the following advantages:
[0034] Without requiring the target person to speak, the presence of a target person in the target direction can be determined based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient. This capability overcomes the limitations of existing technologies by allowing for determination of a target person's presence even when they are silent. Attached Figure Description
[0035] Figure 1A and Figure 1B This schematic diagram illustrates the environment of the method for determining the direction of a person in the environment according to an embodiment of this application.
[0036] Figure 2 A flowchart illustrating a method for determining the direction of a person in an environment according to Embodiment 1 of this application is shown schematically.
[0037] Figure 3 yes Figure 2 Flowchart of step S200;
[0038] Figure 4 yes Figure 2 Flowchart of step S202;
[0039] Figure 5 yes Figure 2 Flowchart of step S204;
[0040] Figure 6 A flowchart illustrating a method for determining the target direction of a target person according to an embodiment of this application is shown schematically.
[0041] Figure 7 A flowchart illustrating a method for determining the target direction of a target person through comprehensive decision-making according to an embodiment of this application is shown in the schematic diagram.
[0042] Figure 8 A flowchart illustrating a method for guiding multiple microphone modules according to an embodiment of this application is shown schematically;
[0043] Figure 9 This is a schematic diagram of an example of a method for determining the direction of a person in an environment according to an embodiment of this application;
[0044] Figure 10 A block diagram of a system for determining the direction of a person in an environment according to Embodiment 2 of this application is shown schematically.
[0045] Figure 11 The illustration shows a schematic diagram of the hardware architecture of a computer device suitable for implementing a method for determining the direction of a person in an environment, according to Embodiment 3 of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0047] It should be noted that the descriptions involving "first," "second," etc., in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0048] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of the steps, but are only used to facilitate the description of this application and to distinguish each step, and therefore should not be construed as a limitation of this application.
[0049] The following is an explanation of the terms used in this application:
[0050] RT60 is an abbreviation for Reverberation Time 60dB, which refers to the time it takes for energy (audio signal) to decay 60dB from its peak.
[0051] Reverberation factor: In this application, it refers to the reverberation time RT60.
[0052] Figure 1A and Figure 1B The illustration shows an environmental diagram of a method for determining the direction of a person in an environment according to an embodiment of this application.
[0053] In an exemplary embodiment of this application, the environment may be a room or a conference room, and multiple microphone modules are provided in multiple directions within the environment, such as three microphone modules 10, 20, and 30, located in the upper left, right, and lower left of the environment, respectively. For the hearing aid wearer, the three microphone modules 10, 20, and 30 are remote microphone modules.
[0054] As an example, each microphone module 10, 20, 30 includes a speaker (also called a loudspeaker) 10a, 20a, 30a and a microphone 10b, 20b, 30b, respectively. The speakers 10a, 20a, 30a are used to play audio signals, and the microphones 10b, 20b, 30b are used to collect audio signals from the environment.
[0055] The method for determining a person in a target direction in the environment according to the embodiments of this application includes: monitoring multiple reverberation coefficients in multiple directions through the multiple microphone modules 10, 20, and 30; determining the average reverberation coefficient of each direction within a preset time period based on the multiple reverberation coefficients in multiple directions; and determining whether a target person appears in the target direction based on the difference between the target reverberation coefficient of the target direction and the average reverberation coefficient.
[0056] This application embodiment does not require the target person to speak; it can determine whether a target person is present in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient. Even if the target person does not speak, it can still determine whether a target person is present in the target direction, thus overcoming the shortcomings of existing technologies.
[0057] In an exemplary embodiment of this application, three speakers 10a, 20a, and 30a each play an ultrasonic signal at different time intervals. This ultrasonic signal is transmitted to a person and simultaneously reflected back, due to the human body's reflection coefficient. When the corresponding microphones 10b, 20b, and 30b detect the feedback echo, they determine that the target speech is located in that direction. A beam is then formed, pointing in that direction to pick up the target speech.
[0058] Three loudspeakers (10a, 20a, and 30a) form a loudspeaker array, creating a directional loudspeaker. The beam of sound is directed in multiple directions. Since the walls in a room typically have a high reflectivity, it is set to alpha1. Humans, on the other hand, have a relatively low reflectivity, set to alpha2. When the sound signal played by the multi-directional loudspeaker array hits a person, the intensity of the reflected energy is directly reduced.
[0059] Once a target person is identified as appearing in the target direction, a comprehensive decision-making process can be used to determine their direction. Specifically, this involves processing the signal emitted by the loudspeaker, the direction of the feedback echo collected by the microphone, and the direction of the loudspeaker's feedback signal. A comprehensive decision analysis is then performed, combining the beamform from the microphone with the direction of the feedback signal from the loudspeaker. This allows for a better determination of the target person's direction.
[0060] Once the target person's direction is determined, it can guide the selection of the speaker array and the shape of the microphone beam. For example, once the target person is identified, the beam is aimed at that person, thus enabling better pickup of their voice.
[0061] The advantages of this application's embodiments include: beamforming can still be formed even when the target person does not need to speak, or even when the target person speaks very little, to implement an optimized sound processing strategy for the target person, such as directional pickup or directional speaker.
[0062] Several embodiments will be provided below, which can be used to implement the method described above for determining the direction of a person in an environment. For ease of understanding, a microphone module will be used as an example in the following description.
[0063] Example 1
[0064] Figure 2 The flowchart illustrates a method for determining the direction of a person in an environment according to Embodiment 1 of this application.
[0065] like Figure 2As shown, the method for determining a person in a target direction in the environment may include steps S200 to S204, wherein: step S200, monitoring multiple reverberation coefficients in multiple directions through the multiple microphone modules; step S202, determining the average reverberation coefficient of each direction within a preset time period based on the multiple reverberation coefficients in multiple directions; step S204, determining whether a target person appears in the target direction based on the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient.
[0066] pass Figure 2 The steps in this embodiment do not require the target person to speak. They can determine whether a target person is present in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient. Even if the target person does not speak, it is possible to determine whether a target person is present in the target direction, thus overcoming the shortcomings of existing technologies.
[0067] In an exemplary embodiment, such as Figure 3 As shown, step S200 can be implemented through steps S300 to S306: Step S300, playing a first ultrasonic signal through the speaker of each microphone module; Step S302, acquiring a second ultrasonic signal through the microphone of each microphone module, the second ultrasonic signal being the ultrasonic signal formed by the reflection of the first ultrasonic signal from the environment; Step S304, determining the speaker-to-microphone impact response function of each microphone module based on the first ultrasonic signal and the second ultrasonic signal; Step S306, determining the reverberation coefficient in each direction based on the impact response function. As an example, the impact response function can be calculated using adaptive filtering.
[0068] pass Figure 3 The steps involve monitoring multiple reverberation coefficients in each direction through the combination of the speaker and microphone in each microphone module.
[0069] In an exemplary embodiment, such as Figure 4 As shown, step S202 may include step S400, which involves tracking and analyzing the reverberation coefficient of each direction within the preset time period, and determining the average value of the reverberation coefficient of each direction within the preset time period.
[0070] pass Figure 4 The steps involved can track and analyze the reverberation coefficient of each direction over a long period of time to obtain an average value, which serves as a basic quantity. This facilitates subsequent determination of whether a target person is present in the target direction based on the difference between the target reverberation coefficient of the target direction and the average value of the reverberation coefficient.
[0071] In an exemplary embodiment, such as Figure 5As shown, step S204 may include step S500, whereby if the difference between the target reverberation coefficient in the target direction (e.g., the RT60 value in the current direction) and the average reverberation coefficient in the target direction (e.g., the long-term average RT60 value) is greater than a threshold, the target person is determined to be present in the target direction; if the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is not greater than the threshold, the target person is determined not to be present in the target direction.
[0072] For example, if in the direction of speaker 10a, |current RT60 value - long-term average RT60 value| > threshold, then it means that there is a high probability that a target person has appeared in the direction that speaker 10a is pointing, and therefore it is determined that there is someone in that direction. And so on.
[0073] pass Figure 5 The steps do not require the target person to speak; they can determine whether the target person is present in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient.
[0074] Figure 6 The flowchart illustrates a method for determining the target direction of a target person according to an embodiment of this application.
[0075] like Figure 6 As shown, the method for determining the target direction of a target person may include steps S600 to S604, wherein: step S600, a microphone array is formed by multiple microphones of the multiple microphone modules to form multiple beams; step S602, sound signals within the range of multiple beams are picked up, the range of multiple beams corresponds to multiple directions in the environment; step S604, the target direction of the target person is determined according to the range of the beams in which the sound signals are located.
[0076] pass Figure 6 This process can improve the accuracy of the judgment, and at the same time, a microphone is used to form a beam to pick up the sound of the target person (such as a hearing aid wearer) so as to determine the direction of the sound.
[0077] Figure 7 The flowchart illustrates a method for determining the target direction of a target person through comprehensive decision-making according to an embodiment of this application.
[0078] like Figure 7As shown, the method for comprehensively determining the target direction of a target person may include steps S700 to S704, wherein: step S700, the difference between the target reverberation coefficient of the target direction and the average reverberation coefficient of the target direction is used as a first decision factor; step S702, the magnitude of the sound signal within the plurality of beam ranges is used as a second decision factor; step S704, the target direction of the target person is comprehensively determined based on the first decision factor and the second decision factor.
[0079] For example, if the first decision factor is large, then increase its weight or confidence level. If the second decision factor is large, then increase its weight or confidence level. And so on.
[0080] pass Figure 7 The steps, based on the first and second decision factors, to make a comprehensive decision-making judgment can improve the accuracy of judging the target person's target direction.
[0081] Figure 8 A flowchart illustrating a method for guiding multiple microphone modules according to an embodiment of this application is shown schematically.
[0082] like Figure 8 As shown, the method for guiding multiple microphone modules may include step S800, which guides the selection of the speaker array and the shape of the microphone beam for the multiple microphone modules based on the target direction of the target person. For example, when it is determined that the target speaker is coming from direction A, a microphone beam can be formed with the main lobe of the beam aligned with A, thereby effectively extracting the speech signal of the target speaker from direction A, and so on. Similarly, a strategy can be adopted whereby, when it is determined that the target speaker is coming from direction A, a speaker array beam can be formed with the main lobe of the beam aligned with A, thereby improving the sound quality and volume heard by the target person from direction A compared to other directions, thus improving the signal-to-noise ratio of the signal played by the speaker array heard by the target person from direction A. At the same time, the signal strength heard from other directions is relatively low.
[0083] pass Figure 8 The process involves determining the target person's direction, which guides the selection of the speaker array and the shape of the microphone beam. Once the target person is identified, the beam is aimed at that person to better pick up their voice.
[0084] Figure 9 This is a schematic diagram of an example of a method for determining the direction of a person in the environment according to an embodiment of this application.
[0085] In an exemplary embodiment of this application, the reverberation coefficients (RT60 values) in three directions (upper left, right, and lower left) can be monitored by three microphone modules 10, 20, and 30. The reverberation coefficients in each direction are tracked and analyzed within a preset time period to determine the average value of the reverberation coefficients in each direction within the preset time period.
[0086] For example, three loudspeakers 10a, 20a, and 30a each play an ultrasonic signal (the first ultrasonic signal) at different time intervals. This ultrasonic signal is transmitted to a person and reflected back. When loudspeaker 10a plays the ultrasonic signal, microphone 10b picks up the ultrasonic signal (the second ultrasonic signal). When there is no one in the room, or when there is no target person in the direction of the signal played by loudspeaker 10a, the signal picked up by microphone 10b can be used to calculate its impulse response function. From this impulse response function, the RT60 value (reverberation coefficient) of the system can be calculated. This RT60 value can be tracked and analyzed over a long period to obtain an average value (average reverberation coefficient), which serves as a baseline quantity.
[0087] Step S900: Determine whether there is anyone in the room based on the RT60 value.
[0088] As an example, when someone is positioned in the direction of speaker 10a, the signal being played will be blocked or reflected by that person, causing a change in the system response function (impact response function). Consequently, the calculated RT60 value will typically be larger because the person absorbs some of the reflected signal. This changed RT60 value can be compared to the average RT60 value (average reverberation coefficient). If a change is observed, it can be preliminarily determined that a target person is present in that direction. Similarly, this can be extrapolated to other speaker and microphone combinations to determine the probability of a target person appearing in other directions.
[0089] Step S902: Determine the direction of the target person based on the feedback information.
[0090] For example, if in the direction of speaker 10a, |current RT60 value - long-term average RT60 value| > threshold, then it means that there is a high probability that a target person is present in the direction that speaker 10a is pointing, and therefore it is determined that there is someone in that direction. And so on.
[0091] Steps S904 and S906 involve beamforming and determining the direction of the target person.
[0092] As an example, a microphone array composed of three microphones 10b, 20b, and 30b forms multiple beams to pick up sound signals within multiple beam ranges, which correspond to multiple directions in the environment. For instance, based on the formation of multiple beams, if a person is moving in a specific direction, then the sound of footsteps or voices can be detected, and based on the sound of voices, it can be determined that someone is in that direction.
[0093] Step S908: Comprehensive decision-making and judgment to determine the location of the target person.
[0094] As an example, a weighted decision is made based on "determining the location of the target person based on feedback signals" (e.g., using the difference between the target reverberation coefficient and the average reverberation coefficient in the target direction as the first decision factor) and "determining the location of the target person based on beamforming" (e.g., using the sound volume within multiple beam ranges as the second decision factor). If the difference between |current direction RT60 value - long-term average RT60 value| in "determining the location of the target person based on feedback signals" is large (i.e., the first decision factor is large), then the confidence level of that decision factor is increased. And so on.
[0095] Step S910 guides the selection of the speaker array and the shape of the microphone beam.
[0096] Example 2
[0097] like Figure 10 The diagram illustrates a block diagram of a system 100 for determining the direction of a person in an environment according to Embodiment 2 of this application. The system 100 can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiment of this application. The program module referred to in this embodiment is a series of computer program instruction segments capable of performing a specific function. The following description will specifically introduce the functions of each program module in this embodiment. Specifically, the system 100 for determining the direction of a person in an environment includes the following modules:
[0098] The monitoring module 110 is used to monitor multiple reverberation coefficients in the multiple directions through the multiple microphone modules;
[0099] The determining module 120 is used to determine the average value of the reverberation coefficients in each direction within a preset time period based on the multiple reverberation coefficients in the multiple directions.
[0100] The discrimination module 130 is used to determine whether a target person appears in the target direction based on the difference between the target reverberation coefficient and the average value of the reverberation coefficient in the target direction.
[0101] As an optional embodiment, the monitoring module 110 is further configured to:
[0102] The first ultrasonic signal is played through the speaker of each microphone module;
[0103] The second ultrasonic signal is acquired by the microphone of each microphone module. The second ultrasonic signal is the ultrasonic signal formed by the reflection of the first ultrasonic signal by the environment.
[0104] The speaker-to-microphone impulse response function of each microphone module is determined based on the first ultrasonic segment signal and the second ultrasonic segment signal.
[0105] Based on the impact response function, the reverberation coefficient in each direction is determined.
[0106] As an optional embodiment, the determining module 120 is further configured to:
[0107] The reverberation coefficient in each direction is tracked and analyzed within the preset time period, and the average reverberation coefficient in each direction is determined within the preset time period.
[0108] As an optional embodiment, the discrimination module 130 is further configured to:
[0109] If the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is greater than a threshold, it is determined that the target person appears in the target direction.
[0110] If the difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is not greater than a threshold, it is determined that the target person does not appear in the target direction.
[0111] As an optional embodiment, the discrimination module 130 is further configured to:
[0112] The microphones of the multiple microphone modules are combined to form a microphone array, thereby creating multiple beams;
[0113] The system picks up sound signals within multiple beam ranges, the multiple beam ranges corresponding to multiple directions in the environment;
[0114] The target direction of the target person is determined based on the beam range of the sound signal.
[0115] As an optional embodiment, the discrimination module 130 is further configured to:
[0116] The difference between the target reverberation coefficient in the target direction and the average reverberation coefficient in the target direction is used as the first decision factor;
[0117] The magnitude of the sound signal within the multiple beam ranges is used as the second decision factor;
[0118] Based on the first decision factor and the second decision factor, a comprehensive decision is made to determine the target person's target direction.
[0119] As an optional embodiment, the discrimination module 130 is further configured to:
[0120] Based on the target person's target direction, the selection of the speaker array and the shape of the microphone beam of the multiple microphone modules are guided.
[0121] Example 3
[0122] like Figure 11 The diagram illustrates a hardware architecture schematic of a computer device 10000, according to Embodiment 3 of this application, suitable for implementing a method for determining the direction of a person in an environment. In this embodiment, the computer device 10000 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Figure 11 As shown, the computer device 10000 includes, but is not limited to, at least the following: a memory 10010, a processor 10020, and a network interface 10030 that can communicate and be linked to each other via a system bus. Wherein:
[0123] The memory 10010 includes at least one type of computer-readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of a computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is typically used to store the operating system and various application software installed on the computer device 10000, such as program code for methods to determine the direction of a person in the environment. In addition, the memory 10010 can also be used to temporarily store various types of data that have been output or will be output.
[0124] In some embodiments, processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 10020 is typically used to control the overall operation of computer device 10000, such as performing control and processing related to data interaction or communication with computer device 10000. In this embodiment, processor 10020 is used to run program code stored in memory 10010 or process data.
[0125] Network interface 10030 may include a wireless network interface or a wired network interface, which is typically used to establish a communication link between computer device 10000 and other computer devices. For example, network interface 10030 is used to connect computer device 10000 to an external terminal via a network, establishing a data transmission channel and communication link between computer device 10000 and the external terminal. The network can be an intranet, the Internet, Global System for Mobile Communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, or other wireless or wired networks.
[0126] It should be pointed out that, Figure 11 Only computer devices with components 10010-10030 are shown; however, it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead.
[0127] In this embodiment, the method for determining the direction of a person in the environment, stored in the memory 10010, can be further divided into one or more program modules and executed by one or more processors (processor 10020 in this embodiment) to complete the embodiment of this application.
[0128] Example 4
[0129] This application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method for determining the direction of a person in an environment as described in the embodiments.
[0130] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the computer-readable storage medium can also include both internal storage units and external storage devices of the computer device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method for determining the direction of a person in the environment. Furthermore, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or will be output.
[0131] Obviously, those skilled in the art should understand that the modules or steps of the embodiments of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of this application are not limited to any particular combination of hardware and software.
[0132] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method of discriminating a target person in an environment, characterized by, A plurality of microphone modules are arranged in a plurality of directions of the environment, and the method for identifying a target person in a target direction in the environment comprises: Monitoring a plurality of reverberation degree coefficients of the plurality of directions by the plurality of microphone modules; Determining a mean value of the reverberation degree coefficients of each direction in a preset time period according to the plurality of reverberation degree coefficients of the plurality of directions; Identifying whether a target person appears in the target direction according to a difference between a target reverberation degree coefficient of the target direction and the mean value of the reverberation degree coefficients; Wherein, identifying whether a target person appears in the target direction according to a difference between a target reverberation degree coefficient of the target direction and the mean value of the reverberation degree coefficients comprises: in the case that the difference between the target reverberation degree coefficient of the target direction and the mean value of the reverberation degree coefficients of the target direction is greater than a threshold value, identifying that the target person appears in the target direction; in the case that the difference between the target reverberation degree coefficient of the target direction and the mean value of the reverberation degree coefficients of the target direction is not greater than a threshold value, identifying that the target person does not appear in the target direction.
2. The method of discriminating a person in a direction of an object in an environment according to claim 1, characterized by, Each microphone module comprises a loudspeaker and a microphone; The monitoring of the plurality of reverberation degree coefficients of the plurality of directions by the plurality of microphone modules comprises: Playing a first ultrasonic segment signal through the loudspeaker of each microphone module; Collecting a second ultrasonic segment signal through the microphone of each microphone module, the second ultrasonic segment signal being an ultrasonic segment signal formed by the reflection of the first ultrasonic segment signal in the environment; Determining an impulse response function from the loudspeaker to the microphone of each microphone module according to the first ultrasonic segment signal and the second ultrasonic segment signal; Determining the reverberation degree coefficient of each direction according to the impulse response function.
3. The method of identifying a person of interest in an environment of claim 1, wherein, The determining of the mean value of the reverberation degree coefficients of each direction in a preset time period according to the plurality of reverberation degree coefficients of the plurality of directions comprises: Tracking and analyzing the reverberation degree coefficients of each direction in the preset time period to determine the mean value of the reverberation degree coefficients of each direction in the preset time period.
4. The method of discriminating a person in a direction of a target in an environment according to any one of claims 1 to 3, characterized by, Further comprising: Forming a plurality of beams by a plurality of microphones of the plurality of microphone modules; Picking up sound signals in a plurality of beam ranges, the plurality of beam ranges corresponding to a plurality of directions in the environment; Determining the target direction of the target person according to the beam range in which the sound signal is located.
5. The method of discriminating a person directing a target in an environment according to claim 4, wherein, Further comprising: Taking the difference between the target reverberation degree coefficient of the target direction and the mean value of the reverberation degree coefficients of the target direction as a first decision factor; Taking the size of the sound signal in the plurality of beam ranges as a second decision factor; Comprehensively determining the target direction of the target person according to the first decision factor and the second decision factor; Wherein, comprehensively determining the target direction of the target person according to the first decision factor and the second decision factor comprises: if the first decision factor is large, then the weight or confidence of the first decision factor is increased; if the second decision factor is large, then the weight or confidence of the second decision factor is increased.
6. The method of discriminating a person directing an object in an environment according to claim 5, wherein, Further comprising: According to the target direction of the target person, guiding the array selection of the loudspeakers and the shape of the microphone beams of the plurality of microphone modules.
7. A system for discriminating a direction of a target person in an environment, characterized by, A plurality of microphone modules are arranged in a plurality of directions of the environment, and the system for identifying a target person in a target direction in the environment comprises: a monitoring module configured to monitor a plurality of reverberation degree coefficients of the plurality of directions by the plurality of microphone modules; a determining module configured to determine an average value of the reverberation degree coefficients of each direction in a preset time period according to the plurality of reverberation degree coefficients of the plurality of directions; an identifying module configured to identify whether a target person appears in the target direction according to a difference between a target reverberation degree coefficient of the target direction and the average value of the reverberation degree coefficients. The identifying whether the target person appears in the target direction according to the difference between the target reverberation degree coefficient of the target direction and the average value of the reverberation degree coefficients comprises: identifying that the target person appears in the target direction when the difference between the target reverberation degree coefficient of the target direction and the average value of the reverberation degree coefficients of the target direction is greater than a threshold value; and identifying that the target person does not appear in the target direction when the difference between the target reverberation degree coefficient of the target direction and the average value of the reverberation degree coefficients of the target direction is not greater than the threshold value.
8. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor is configured to execute the computer program to implement the steps of the method for identifying a target person in a target direction in an environment according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executable by at least one processor to enable the at least one processor to execute the steps of the method for identifying a target person in a target direction in an environment according to any one of claims 1 to 6.
Citation Information
Patent Citations
Beam information arrival direction determination method and device, and storage medium
CN112558004A
Method for detecting, positioning and identifying characteristics in pipeline by using intelligent acoustic technology
CN114721037A