Sound source direction determination method, device, terminal, storage medium and product
By determining the similarity of voice signals in the pickup direction on the terminal and combining it with image recognition, the problem of inaccurate sound source direction in a noisy environment is solved, and the sound source direction can be accurately determined in a noisy environment and the clarity of voice signal collection is improved.
Patent Information
- Application Number
- CN202210558040.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-05-19
AI Technical Summary
In the presence of noise in the surrounding environment, the accuracy of the sound source direction determined by the terminal is poor.
By determining the similarity of the voice signal in the pickup direction, if it is greater than the preset threshold, the camera is controlled to rotate and collect images. The direction of the target sound source is determined by combining the image object recognition results and the camera orientation.
The accuracy of the sound source direction is improved, ensuring that the sound source direction can be accurately determined even in a noisy environment, and improving the clarity of voice signal acquisition and interaction efficiency.
Smart Images

Figure CN115035187B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sound source localization, and in particular to a method, device, terminal, storage medium and product for determining the direction of a sound source. Background Art
[0002] Currently, some terminals can recognize user voice signals and thus interact with the user. However, in order to save power, the terminal is generally in a dormant state and wakes up only when a wake-up command is received.
[0003] In related technologies, to improve the clarity of collected voice signals, the terminal can use the voice signal corresponding to the wake-up command to determine the direction of the sound source, and then collect voice signals based on this sound source direction. However, in the presence of ambient noise, the accuracy of the determined sound source direction is poor. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, terminal, storage medium, and product for determining the direction of a sound source, which can improve the accuracy of the direction of the sound source. The technical solution is as follows:
[0005] According to one aspect of an embodiment of the present application, a method for determining a sound source direction is provided, the method comprising:
[0006] Determining a speech signal of the collected target speech signal in at least one sound pickup direction;
[0007] If there is a similarity greater than a preset similarity threshold among the similarities corresponding to the at least one sound pickup direction, determining the initial sound source direction, wherein the similarity is the similarity between the voice signal corresponding to the sound pickup direction and the preset wake-up word;
[0008] If the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, the camera on the terminal is controlled to rotate and capture an image, and the target sound source direction is determined based on the object recognition result of the currently captured image and the current orientation of the camera. The target sound pickup direction is the direction corresponding to the maximum similarity among the at least one sound pickup direction.
[0009] In one possible implementation, the object recognition result indicates whether the object is not recognized or recognized in the captured image, and determining the target sound source direction based on the object recognition result of the currently captured image and the current orientation of the camera includes:
[0010] If the camera has not passed through the first direction and an object is recognized in the currently captured image, recording the current direction of the camera and controlling the camera to continue rotating and capturing images;
[0011] If the camera has passed the first direction but has not yet reached the second direction, and an object is recognized in the currently captured image, the current direction of the camera is determined as the target sound source direction, and the camera is controlled to stop rotating; or, if the camera is currently facing the second direction, no object is recognized in the currently captured image, and the direction has been recorded, the target sound source direction is determined based on the recorded direction, and the camera is controlled to rotate back to the target sound source direction;
[0012] The first direction is the first direction that the camera passes through between the initial sound source direction and the target sound pickup direction, and the second direction is the second direction that the camera passes through between the initial sound source direction and the target sound pickup direction.
[0013] In a possible implementation, the method further includes:
[0014] If the current direction of the camera is the second direction, no object is recognized in the currently captured image, and the direction is not recorded, the second direction is determined as the target sound source direction, and the camera is controlled to stop rotating.
[0015] In a possible implementation, determining the target sound source direction based on the recorded orientation includes:
[0016] If the number of recorded directions is 1, the recorded direction is determined as the target sound source direction;
[0017] If there are multiple recorded directions, the direction with the smallest angle with the first direction among the multiple recorded directions is determined as the target sound source direction.
[0018] In a possible implementation, before controlling the camera on the terminal to rotate and capture images, the method further includes:
[0019] Determine the direction parameters corresponding to the counterclockwise rotation direction and the direction parameters corresponding to the clockwise rotation direction respectively;
[0020] determining a target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter;
[0021] The controlling the rotation of the camera on the terminal includes:
[0022] Control the camera on the terminal to rotate according to the target rotation direction.
[0023] In a possible implementation, respectively determining the direction parameter corresponding to the counterclockwise rotation direction and the direction parameter corresponding to the clockwise rotation direction includes:
[0024] If the number of the at least one sound pickup direction is greater than 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determine at least one intermediate direction, and determine a weighted average of the similarities corresponding to the at least one intermediate direction as the direction parameter corresponding to the rotation direction, the intermediate direction being a sound pickup direction that is intermediate between the current orientation of the camera and the target sound pickup direction in the rotation direction;
[0025] Determining the target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter includes: determining the rotation direction corresponding to the determined maximum direction parameter as the target rotation direction.
[0026] In a possible implementation, respectively determining the direction parameter corresponding to the counterclockwise rotation direction and the direction parameter corresponding to the clockwise rotation direction includes:
[0027] If the number of the at least one sound pickup direction is less than or equal to 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determining a first angle between the current orientation of the camera and the initial sound source direction, and a second angle between the current orientation of the camera and the target sound pickup direction in the rotation direction, and determining the maximum angle between the first angle and the second angle as the direction parameter corresponding to the rotation direction;
[0028] The determining of the target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter includes: determining the rotation direction corresponding to the determined minimum direction parameter as the target rotation direction.
[0029] In a possible implementation, the method further includes:
[0030] If the angle between the initial sound source direction and the target sound pickup direction is less than or equal to the preset angle threshold, the initial sound source direction is determined as the target sound source direction.
[0031] In a possible implementation, the target voice signal includes an initial voice signal corresponding to each of the sound pickup directions, and determining a voice signal of the collected target voice signal in at least one sound pickup direction includes:
[0032] For each sound pickup direction, noise suppression is performed on the initial speech signals corresponding to the other sound pickup directions in the target speech signal except the sound pickup direction, and the target speech signal after noise suppression is determined as the speech signal corresponding to the sound pickup direction.
[0033] In a possible implementation, the method further includes:
[0034] The voice signal corresponding to the at least one sound pickup direction is input into a preset wake-up model to obtain the similarity corresponding to the at least one sound pickup direction. The preset wake-up model is used to determine the similarity between the input voice signal and the preset wake-up word.
[0035] According to another aspect of an embodiment of the present application, a device for determining a sound source direction is provided, the device comprising:
[0036] A signal determination module, configured to determine a speech signal of the collected target speech signal in at least one sound pickup direction;
[0037] a first direction determination module, configured to determine an initial sound source direction if a similarity greater than a preset similarity threshold exists among the similarities corresponding to the at least one sound pickup direction, wherein the similarity is a similarity between the voice signal corresponding to the sound pickup direction and a preset wake-up word;
[0038] The second direction determination module is used to control the camera on the terminal to rotate and capture images if the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, and determine the target sound source direction based on the object recognition result of the currently captured image and the current orientation of the camera. The target sound pickup direction is the direction corresponding to the maximum similarity among the at least one sound pickup direction.
[0039] In a possible implementation, the object recognition result indicates that the object is not recognized or is recognized in the captured image, and the second direction determination module is configured to:
[0040] If the camera has not passed through the first direction and an object is recognized in the currently captured image, recording the current direction of the camera and controlling the camera to continue rotating and capturing images;
[0041] If the camera has passed the first direction but has not yet reached the second direction, and an object is recognized in the currently captured image, the current direction of the camera is determined as the target sound source direction, and the camera is controlled to stop rotating; or, if the camera is currently facing the second direction, no object is recognized in the currently captured image, and the direction has been recorded, the target sound source direction is determined based on the recorded direction, and the camera is controlled to rotate back to the target sound source direction;
[0042] The first direction is the first direction that the camera passes through between the initial sound source direction and the target sound pickup direction, and the second direction is the second direction that the camera passes through between the initial sound source direction and the target sound pickup direction.
[0043] In a possible implementation, the apparatus further includes:
[0044] The second direction determination module is further used to determine the second direction as the target sound source direction and control the camera to stop rotating if the current direction of the camera is the second direction, no object is recognized in the currently captured image, and the direction is not recorded.
[0045] In a possible implementation, the second direction determining module is configured to:
[0046] If the number of recorded directions is 1, the recorded direction is determined as the target sound source direction;
[0047] If there are multiple recorded directions, the direction with the smallest angle with the first direction among the multiple recorded directions is determined as the target sound source direction.
[0048] In a possible implementation, the apparatus further includes:
[0049] a rotation direction determination module, configured to determine a direction parameter corresponding to a counterclockwise rotation direction and a direction parameter corresponding to a clockwise rotation direction, respectively; and determine a target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameters;
[0050] The second direction determining module is configured to:
[0051] Control the camera on the terminal to rotate according to the target rotation direction.
[0052] In a possible implementation, the rotation direction determination module is configured to:
[0053] If the number of the at least one sound pickup direction is greater than 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determine at least one intermediate direction, and determine a weighted average of the similarities corresponding to the at least one intermediate direction as the direction parameter corresponding to the rotation direction, the intermediate direction being a sound pickup direction that is intermediate between the current orientation of the camera and the target sound pickup direction in the rotation direction;
[0054] The rotation direction corresponding to the determined maximum direction parameter is determined as the target rotation direction.
[0055] In a possible implementation, the rotation direction determination module is configured to:
[0056] If the number of the at least one sound pickup direction is less than or equal to 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determining a first angle between the current orientation of the camera and the initial sound source direction, and a second angle between the current orientation of the camera and the target sound pickup direction in the rotation direction, and determining the maximum angle between the first angle and the second angle as the direction parameter corresponding to the rotation direction;
[0057] The rotation direction corresponding to the determined minimum direction parameter is determined as the target rotation direction.
[0058] In a possible implementation, the apparatus further includes:
[0059] The second direction determining module is configured to determine the initial sound source direction as the target sound source direction if the angle between the initial sound source direction and the target sound pickup direction is less than or equal to the preset angle threshold.
[0060] In one possible implementation, the target speech signal includes an initial speech signal corresponding to each of the pickup directions, and the signal determination module is used to perform noise suppression on the initial speech signals corresponding to other pickup directions in the target speech signal except the pickup direction for each of the pickup directions, and determine the target speech signal after noise suppression as the speech signal corresponding to the pickup direction.
[0061] In a possible implementation, the apparatus further includes:
[0062] A similarity determination module is used to input the voice signal corresponding to the at least one pickup direction into a preset wake-up model to obtain the similarity corresponding to the at least one pickup direction, and the preset wake-up model is used to determine the similarity between the input voice signal and the preset wake-up word.
[0063] According to another aspect of an embodiment of the present application, a terminal is provided, comprising a processor and a memory, wherein the memory stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the sound source direction determination method as described in any one of the possible implementation methods described above.
[0064] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the sound source direction determination method described in any of the above possible implementation methods.
[0065] According to another aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program code, wherein the computer program code is stored in a computer-readable storage medium, a processor reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to implement the method for determining the direction of a sound source as described in any one of the possible implementations described above.
[0066] An embodiment of the present application provides a sound source direction determination solution. If the similarity between the voice signal corresponding to a certain pickup direction and a preset wake-up word is greater than a preset similarity threshold, it indicates that the voice signal is likely to be a voice signal used to wake up the terminal. Then the initial sound source direction can be determined. If the angle between the initial sound source direction and the target pickup direction with the greatest similarity is large, it indicates that there may be noise in the environment where the terminal is located. Then the accuracy of the initial sound source direction and the target pickup direction is not high. At this time, the target sound source direction can be determined in combination with the object recognition result of the collected image, so that the accuracy of the determined target sound source direction is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0068] Figure 1 This is a flow chart of a method for determining a sound source direction provided by an embodiment of the present application;
[0069] Figure 2 This is a flowchart of another method for determining the direction of a sound source provided by an embodiment of the present application;
[0070] Figure 3 Schematic diagram of a sound source direction determination process provided by an embodiment of the present application;
[0071] Figure 4 This is a structural block diagram of a device for determining a sound source direction provided by an embodiment of the present application;
[0072] Figure 5 This is a structural block diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0074] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another.
[0075] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0076] It should be noted that the signals (including but not limited to voice signals) and data (including but not limited to data used for processing, storage, and display) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the voice signals and images involved in this application were obtained with full authorization.
[0077] The present embodiment provides a method for determining the direction of a sound source, which is performed by a terminal. Optionally, the terminal includes, but is not limited to, an intelligent robot, a smart home device, a smart wearable device, a smartphone, a tablet computer, a laptop computer, or a desktop computer. Among them, the smart home device includes a voice assistant, a smart TV, a smart mirror, a smart refrigerator, or other smart home devices.
[0078] Optionally, the terminal includes or is connected to a voice collection component for collecting voice signals. The voice collection component may include a microphone array, headphones, or a microphone. Optionally, the terminal also includes or is connected to a camera for collecting images. The camera may be rotatable, for example, relative to other components of the terminal.
[0079] In an embodiment of the present application, the terminal has a voice interaction function. The terminal collects a voice signal based on a voice collection component. If the voice signal is a voice signal containing a preset wake-up word, the terminal is woken up, and then interacts with the target sound source that sends the voice signal. Among them, the target sound source is a user who sends a voice signal containing a preset wake-up word, and the user has the intention to interact with the terminal by voice. In addition to the target sound source, there may be other sound sources in the environment where the terminal is located. Other sound sources can be regarded as interference sound sources. The voice signals emitted by the interference sound sources are irrelevant to the voice interaction and can be regarded as noise. When the target sound source emits a voice signal containing a preset wake-up word, the interference sound source may emit noise, causing the voice signal collected by the terminal to contain both the voice signal emitted by the target sound source and noise. Therefore, the terminal needs to determine the sound source direction corresponding to the target sound source, and subsequently collect voice signals for the determined sound source direction. By identifying the collected voice signals and performing corresponding operations, voice interaction with the target sound source is achieved.
[0080] The sound source direction determination method provided in the embodiment of the present application can be applied to a variety of voice interaction scenarios. The application scenarios of the sound source direction determination method are introduced below. For example, the terminal has a music playback function. If the user wants to control the terminal to play music, the user can first say the preset wake-up word "XXXX". The terminal collects a voice signal containing the preset wake-up word. The sound source direction determination method provided in the embodiment of the present application wakes up the terminal and determines the direction of the sound source. Even if there is noise in the surrounding environment, a more accurate sound source direction can be determined. After the user says "play music", the terminal can collect voice signals based on the determined sound source direction, thereby collecting more accurate voice signals, recognizing the voice signals, and performing corresponding operations, such as playing music, thereby improving the interaction efficiency with the user and improving the user experience.
[0081] It should be noted that the above application scenarios are only exemplary descriptions and do not limit the voice interaction scenarios. In addition to being applied to the above scenarios, this application can also be applied to any other voice interaction scenarios.
[0082] Figure 1 This is a flow chart of a method for determining the direction of a sound source provided by an embodiment of the present application. The method is executed by the terminal, see Figure 1 , the method comprises the following steps:
[0083] 101. Determine a speech signal of a collected target speech signal in at least one sound pickup direction.
[0084] The target voice signal is any voice signal collected by the terminal. Optionally, the terminal is provided with a voice collection component for collecting voice signals. Accordingly, the terminal collects the target voice signal through the voice collection component. After collecting the target voice signal, the terminal determines the sound source direction corresponding to the target voice signal based on the target voice signal, so that it can subsequently collect voice signals based on the sound source direction to improve the accuracy of the collected voice signal. The sound source direction corresponding to the voice signal is the direction of the sound source emitting the voice signal relative to the terminal.
[0085] The pickup direction is the direction in which the voice signal is picked up, and the terminal is pre-set with at least one pickup direction. Optionally, when the number of at least one pickup direction is multiple, there is an angle between two adjacent pickup directions, and the angle can be the same or different. For example, the four pickup directions include a 30-degree pickup direction, a 60-degree pickup direction, a 90-degree pickup direction, and a 120-degree pickup direction. In this case, the angle between two adjacent pickup directions is the same, all 30 degrees. For another example, the three pickup directions include a 45-degree pickup direction, a 90-degree pickup direction, and a 180-degree pickup direction. In this case, the angle between any two adjacent pickup directions is different. The embodiment of the present application does not limit the size of the angle between two adjacent pickup directions.
[0086] This embodiment of the present application extracts the voice signal corresponding to each pickup direction from the target voice signal. Each pickup direction may be the direction of the sound source, so a pickup direction can be selected from at least one pickup direction and determined as the sound source direction. However, since the pickup directions are pre-set, and the sound source emitting the target voice signal may be in any direction relative to the terminal, there may be a deviation between the selected pickup direction and the actual sound source direction.
[0087] 102. If there is a similarity greater than a preset similarity threshold among the similarities corresponding to at least one sound pickup direction, then determine the initial sound source direction, where the similarity is the similarity between the voice signal corresponding to the sound pickup direction and the preset wake-up word.
[0088] Among them, the preset similarity threshold can be set as needed, and the embodiment of the present application does not limit this. For example, the preset similarity threshold is a real number between 0 and 1, such as 0.7, 0.8 or 0.9. The preset wake-up word is used to wake up the terminal. The preset wake-up word can be set as needed, and the embodiment of the present application does not limit this. For example, the preset wake-up word is "Xiaobai" or "Mirror Mirror".
[0089] To save power consumption, the terminal is generally in a dormant state, and the terminal can also collect voice signals in the dormant state. In an embodiment of the present application, the terminal determines the similarity between the voice signal corresponding to each pickup direction and the preset wake-up word, and determines whether to wake up the terminal based on the size of the similarity corresponding to at least one pickup direction. If the similarity between the voice signal and the preset wake-up word is greater than the preset similarity threshold, it means that the voice signal and the preset wake-up word are relatively similar, then the voice signal is likely to contain the preset wake-up word and can be used to wake up the terminal; if the similarity between the voice signal and the preset wake-up word is less than or equal to the preset similarity threshold, it means that the voice signal and the preset wake-up word are not very similar, then the voice signal is likely not a voice signal used to wake up the terminal.
[0090] Accordingly, if a similarity greater than a preset similarity threshold exists among the similarities corresponding to at least one sound pickup direction, the terminal may be awakened. If a similarity greater than the preset similarity threshold does not exist among the similarities corresponding to at least one sound pickup direction, it indicates that the target voice signal may not be a voice signal for waking up the terminal. In this case, the terminal may not be awakened and may remain in a dormant state. While in the dormant state, if a voice signal from the surrounding environment is collected, step 101 is executed.
[0091] It should be noted that, in embodiments of the present application, upon determining that the similarity corresponding to at least one sound pickup direction is greater than a preset similarity threshold, the terminal may be awakened, and subsequent steps, such as determining the initial sound source direction, may be performed. Alternatively, embodiments of the present application may awaken the terminal after determining the target sound source direction, which is not limited in the embodiments of the present application.
[0092] When a similarity greater than a preset similarity threshold exists among the similarities corresponding to at least one pickup direction, the terminal determines the initial sound source direction. Optionally, the terminal determines the initial sound source direction based on the target voice signal through a DOA (Direction Of Arrival) estimation algorithm or other sound source localization algorithm. However, when there is noise in the surrounding environment, the collected target voice signal is likely to also contain noise, and the determined initial sound source direction may also be affected by the noise, resulting in the determined initial sound source direction being inaccurate.
[0093] The greater the similarity corresponding to the sound pickup direction, the more likely the sound pickup direction is the direction of the sound source. Accordingly, the terminal determines the direction corresponding to the maximum similarity in at least one sound pickup direction as the target sound pickup direction. Then, the terminal performs the following step 103 to determine the target sound source direction based on the initial sound source direction and the target sound pickup direction, and determines the target sound source direction as the final sound source direction, thereby collecting the voice signal based on the target sound source direction.
[0094] 103. If the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, the camera on the control terminal is rotated and an image is captured. The target sound source direction is determined based on the object recognition result of the currently captured image and the current orientation of the camera. The target sound pickup direction is the direction corresponding to the maximum similarity among at least one sound pickup direction.
[0095] The preset angle threshold can be set as needed, and is not limited in this embodiment of the present application. Optionally, if there are multiple sound pickup directions and the angle between two adjacent sound pickup directions is the same, the preset angle threshold can be half the angle between the two adjacent sound pickup directions. If there are multiple sound pickup directions and the angle between two adjacent sound pickup directions is different, or if there is only one sound pickup direction, the preset angle threshold can be an angle threshold set as needed.
[0096] If the angle between the initial sound source direction and the target pickup direction is greater than the preset angle threshold, it means that the initial sound source direction is quite different from the target pickup direction, and the accuracy of the initial sound source direction and the target pickup direction is low. When the object wakes up the terminal, the object is likely to be close to the terminal. By controlling the camera to rotate and capture the image, the target sound source direction can be determined by combining the object recognition result of the image and the current orientation of the camera. If the angle between the initial sound source direction and the target pickup direction is less than or equal to the preset angle threshold, it means that the initial sound source direction and the target pickup direction are not much different and are relatively close. The initial sound source direction or the target pickup direction is likely to be the true sound source direction, and the terminal can directly determine the initial sound source direction or the target pickup direction as the target sound source direction.
[0097] An embodiment of the present application provides a sound source direction determination solution. If the similarity between the voice signal corresponding to a certain pickup direction and the preset wake-up word is greater than a preset similarity threshold, it indicates that the voice signal is likely to be a voice signal used to wake up the terminal. Then the initial sound source direction can be determined. If the angle between the initial sound source direction and the target pickup direction with the greatest similarity is large, it indicates that there may be noise in the environment where the terminal is located. Then the accuracy of the initial sound source direction and the target pickup direction is not high. At this time, the target sound source direction can be determined in combination with the object recognition result of the collected image, so that the accuracy of the determined target sound source direction is higher.
[0098] Figure 2 This is a flow chart of another method for determining the direction of a sound source provided by an embodiment of the present application. The method is executed by the terminal, see Figure 2 , the method comprises the following steps:
[0099] 201. Determine a speech signal of a collected target speech signal in at least one sound pickup direction.
[0100] The target voice signal is any voice signal collected by the terminal. Optionally, the terminal is provided with a voice collection component for collecting voice signals. Accordingly, the terminal collects the target voice signal through the voice collection component. After collecting the target voice signal, the terminal determines the sound source direction corresponding to the target voice signal based on the target voice signal, so that it can subsequently collect voice signals based on the sound source direction to improve the accuracy of the collected voice signal. The sound source direction corresponding to the voice signal is the direction of the sound source emitting the voice signal relative to the terminal.
[0101] Among them, the pickup direction is the direction of picking up voice signals, and the terminal sets at least one pickup direction in advance. Optionally, when the number of at least one pickup direction is multiple, there is an angle between two adjacent pickup directions, and the angle can be the same or different. For example, the four pickup directions include a 30-degree pickup direction, a 60-degree pickup direction, a 90-degree pickup direction, and a 120-degree pickup direction. In this case, the angle between two adjacent pickup directions is the same, all 30 degrees. For another example, the three pickup directions include a 45-degree pickup direction, a 90-degree pickup direction, and a 180-degree pickup direction. In this case, the angle between any two adjacent pickup directions is different, some angles are 45 degrees, and some angles are 90 degrees. The embodiment of the present application does not limit the size of the angle between two adjacent pickup directions.
[0102] The target speech signal includes an initial speech signal corresponding to each of at least one sound pickup direction. Optionally, step 201 may be implemented by performing noise suppression on the initial speech signals corresponding to sound pickup directions other than the sound pickup direction in the target speech signal for each sound pickup direction, and determining the target speech signal after noise suppression as the speech signal corresponding to the sound pickup direction.
[0103] In the embodiment of the present application, each pickup direction may be a real sound source direction. Then, for each pickup direction, if the pickup direction is the sound source direction, the initial voice signal corresponding to the pickup direction is the voice signal emitted by the sound source, and the initial voice signals corresponding to other pickup directions except the pickup direction are likely to be noise in the environment. By performing noise suppression on the initial voice signals in other pickup directions, the voice signal in the pickup direction is more prominent in the target voice signal after noise suppression and is the main voice signal, thereby being able to more accurately reflect the situation of the voice signal in the pickup direction.
[0104] 202. If there is a similarity greater than a preset similarity threshold among the similarities corresponding to at least one sound pickup direction, then wake up the terminal, where the similarity is the similarity between the voice signal corresponding to the sound pickup direction and a preset wake-up word.
[0105] Among them, the preset similarity threshold can be set as needed, and the embodiment of the present application does not limit this. For example, the preset similarity threshold is a real number between 0 and 1, such as 0.7, 0.8 or 0.9. The preset wake-up word is used to wake up the terminal, and the preset wake-up word can be set as needed, and the embodiment of the present application does not limit this. In order to save power consumption, the terminal is generally in a dormant state, and the terminal can also collect voice signals in a dormant state. After obtaining a voice signal corresponding to at least one pickup direction, determine the similarity between the voice signal corresponding to each pickup direction and the preset wake-up word, and determine whether to wake up the terminal based on the size of the similarity corresponding to at least one pickup direction.
[0106] In one possible implementation, the similarity is determined based on a preset wake-up model, which is used to determine the similarity between the input voice signal and the preset wake-up word. Accordingly, the process of determining the similarity includes: inputting a voice signal corresponding to at least one pickup direction into the preset wake-up model to obtain the similarity corresponding to at least one pickup direction. The preset wake-up model is trained for the preset wake-up word, and the preset wake-up model includes the preset wake-up word. The input data of the preset wake-up model is the voice signal in the pickup direction, and the output data is the similarity between the voice signal and the preset wake-up word. Optionally, the training process of the preset wake-up model includes: using the label corresponding to the sample voice signal as supervision, training the preset wake-up model based on the sample voice signal, wherein the label corresponding to the sample voice signal indicates whether the sample voice signal is the voice signal corresponding to the preset wake-up word. Accordingly, the preset wake-up model is called to determine the predicted similarity between the sample voice signal and the preset wake-up word, determine the loss value based on the predicted similarity and the label, and train the preset wake-up model based on the loss value. The smaller the difference between the predicted similarity and the label, the smaller the loss value, and the more accurate the preset wake-up model's prediction. The larger the difference between the predicted similarity and the label, the larger the loss value, and the less accurate the preset wake-up model's prediction. The training goal of the preset wake-up model is to minimize the loss value, that is, to ensure that the predicted similarity between the sample speech signal and the preset wake-up word approaches the label corresponding to the sample speech signal.
[0107] Since the preset wake-up model can learn the relationship between a large number of sample voice signals and the preset wake-up words during the training process, after the voice signal is input into the trained preset wake-up model, the preset wake-up model can output the similarity between the voice signal and the preset wake-up word. The accuracy of the determined similarity is guaranteed and the determination efficiency is high.
[0108] In another possible implementation, the similarity is determined based on the voice features corresponding to the voice signal and the wake-up word features corresponding to the preset wake-up word. Accordingly, the process of determining the similarity includes: for each voice signal corresponding to the pickup direction, determining the similarity between the voice features corresponding to the voice signal and the wake-up word features corresponding to the preset wake-up word, for example, the similarity is cosine similarity. By calculating the similarity between the voice features corresponding to the voice signal and the wake-up word features corresponding to the preset wake-up word, the amount of calculation is reduced, thereby improving the efficiency of calculating the similarity.
[0109] In an embodiment of the present application, if the similarity between the voice signal and the preset wake-up word is greater than the preset similarity threshold, it means that the voice signal is relatively similar to the preset wake-up word, then the voice signal is likely to contain the preset wake-up word and can be used to wake up the terminal; if the similarity between the voice signal and the preset wake-up word is less than or equal to the preset similarity threshold, it means that the voice signal is not very similar to the preset wake-up word, then the voice signal is likely not a voice signal used to wake up the terminal. Accordingly, if there is a similarity greater than the preset similarity threshold in the similarities corresponding to at least one sound pickup direction, the terminal can be woken up so that the terminal interacts with the sound source. If there is no similarity greater than the preset similarity threshold in the similarities corresponding to at least one sound pickup direction, it means that the target voice signal may not be a voice signal used to wake up the terminal, then the terminal may not be woken up and the terminal may be kept in a dormant state. When in a dormant state, if a voice signal in the surrounding environment is collected, step 201 is executed.
[0110] It should be noted that the embodiment of the present application is described by waking up the terminal when it is determined that the similarity corresponding to at least one sound pickup direction is greater than a preset similarity threshold. After waking up the terminal, the terminal executes the subsequent step 203. However, the embodiment of the present application can also wake up the terminal after determining the target sound source direction, and the embodiment of the present application is not limited to this.
[0111] 203. Determine an initial sound source direction based on the target speech signal.
[0112] Optionally, the terminal determines the initial sound source direction based on the target voice signal using a DOA estimation algorithm or other sound source localization algorithm. However, if there is noise in the surrounding environment, the collected target voice signal is likely to also contain noise. In this case, the determined initial sound source direction may also correspond to the noise, making the initial sound source direction inaccurate.
[0113] The greater the similarity corresponding to the pickup direction, the more likely the pickup direction is the direction of the sound source. Accordingly, after waking up the terminal, the terminal determines the direction corresponding to the maximum similarity in at least one pickup direction as the target pickup direction, and determines the target sound source direction based on the initial sound source direction and the target sound source direction, and determines the target sound source direction as the final sound source direction, thereby collecting the voice signal based on the target sound source direction. Accordingly, the terminal executes the following step 204 or executes step 205.
[0114] 204. If the angle between the initial sound source direction and the target sound pickup direction is less than or equal to a preset angle threshold, the initial sound source direction is determined as the target sound source direction, and the target sound pickup direction is the direction corresponding to the maximum similarity among the at least one sound pickup direction.
[0115] The preset angle threshold can be set as needed, and is not limited in this embodiment of the present application. Optionally, if there are multiple sound pickup directions and the angle between two adjacent sound pickup directions is the same, the preset angle threshold can be half the angle between the two adjacent sound pickup directions. If there are multiple sound pickup directions and the angle between two adjacent sound pickup directions is different, or if there is only one sound pickup direction, the preset angle threshold can be an angle threshold set as needed.
[0116] If the angle between the initial sound source direction and the target sound pickup direction is less than or equal to a preset angle threshold, it means that the initial sound source direction and the target sound pickup direction are similar and relatively close, and the initial sound source direction or the target sound pickup direction is very likely to be the true sound source direction, that is, the accuracy of the initial sound source direction or the target sound pickup direction is high, and the initial sound source direction or the target sound pickup direction can be directly determined as the target sound source direction, which does not require excessive calculations and saves computing resources. The embodiment of the present application is described by taking the initial sound source direction as the target sound source direction as an example.
[0117] 205. If the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, the camera on the control terminal rotates and captures an image, and the target sound source direction is determined based on the object recognition result of the currently captured image and the current orientation of the camera.
[0118] If the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, this indicates that the initial sound source direction differs significantly from the target sound pickup direction, and the accuracy of the initial sound source direction and the target sound pickup direction is low. Considering that the subject is likely close to the terminal when waking up the terminal, the camera can be controlled to rotate and capture an image. The captured image may include the user, so the target sound source direction can be determined by combining the object recognition results of the image and the current camera orientation.
[0119] Optionally, the terminal controls the camera to rotate, and during rotation, the camera also captures images. The terminal obtains the captured images, recognizes the images, and obtains object recognition results for the images. The object recognition results indicate whether an object is not recognized or is recognized in the captured images. Optionally, the objects include people, animals, robots, and the like. The terminal recognizes the images using an object recognition model. Accordingly, the process of determining the object recognition result for the currently captured image includes: inputting the captured image into the object recognition model to obtain an object recognition result for the image. The object recognition model is configured to perform object recognition on the input image. The object recognition model is a pre-trained model. When the object is a person, the object recognition model may be a face recognition model. If the object recognition model recognizes a face in the image, it indicates that the image contains the object; if it does not recognize a face, it indicates that the image does not contain the object. Because the object recognition model can learn from sample images containing objects during training, after the captured image is input into the trained object recognition model, the object recognition model can determine whether the image contains the object and obtain an object recognition result. The accuracy of the determined object recognition result is guaranteed, and the determination efficiency is high.
[0120] In one possible implementation, the terminal controls the camera to continuously capture images from the moment it begins rotating, and recognizes each captured image. Alternatively, given that the camera captures images frequently and the range of object movement is small between the capture of two adjacent images, after capturing a certain number of images, one image is extracted from these images and recognized, thereby reducing the recognition workload and improving efficiency. In another possible implementation, the terminal controls the camera to capture an image at a certain interval from the moment it begins rotating, and recognizes each captured image. This reduces the frequency of image capture and saves power.
[0121] Optionally, based on the object recognition result of the currently captured image and the current orientation of the camera, determining the direction of the target sound source may be accomplished by:
[0122] If the camera has not passed through the first direction and an object is recognized in the currently captured image, the current camera orientation is recorded and the camera is controlled to continue rotating. Before the camera rotates, its orientation is fixed. This orientation is the starting orientation, and the terminal controls the camera to rotate from the starting orientation. During the rotation process, the camera may pass through the initial sound source direction or the target sound pickup direction. The first direction is the first direction of the initial sound source direction and the target sound pickup direction that the camera passes through. The second direction is the second direction of the initial sound source direction and the target sound pickup direction that the camera passes through.
[0123] The object identified in the image is likely the sound source of the target voice signal. If the camera's current orientation has not yet passed the first direction, and the angle between the current orientation and the initial sound source direction and the target sound pickup direction is relatively large, the camera can be controlled to continue rotating, potentially capturing images containing the object during subsequent rotations. The current orientation can also be recorded for later reference.
[0124] During continued rotation, if the camera has passed the first direction but has not yet reached the second direction, and an object is recognized in the currently captured image, the camera's current orientation is determined as the target sound source direction, and the camera is controlled to stop rotating. In this case, the angles between the camera's current orientation and both the initial sound source direction and the target sound pickup direction are relatively small. Since the initial sound source direction and the target sound pickup direction are directions determined in the above process and are relatively close to the true sound source direction, the true sound source direction is likely to be located between the initial sound source direction and the target sound pickup direction. Since the current orientation is very close to the true sound source direction, the current orientation can be considered the target sound source direction. Accordingly, after determining the target sound source direction, the camera can be controlled to stop rotating and stop capturing images, so that the camera can capture voice signals in the target sound source direction.
[0125] Alternatively, if the camera does not recognize an object in the captured image while rotating from a first orientation to a second orientation, the camera is controlled to continue rotating and capturing images until it reaches the second orientation. Accordingly, if the camera is currently facing the second orientation, no object is recognized in the currently captured image, and the orientation has been recorded, the target sound source direction is determined based on the recorded orientation, and the camera is controlled to rotate toward the target sound source direction. In this case, if the object recognition result indicates that the object was not recognized, indicating that the image does not contain the object, that is, the object is not currently within the camera's capture range, then the current orientation is likely not the true sound source direction, and the camera can be controlled to continue rotating. However, if the camera rotates to the second orientation but no image containing the object is captured in either the first or second orientations, the target sound source direction can be determined based on whether the orientation has been recorded. If the orientation has been recorded, indicating that an image containing the object was captured between the starting orientation and the first orientation, then, given that the object is likely the sound source emitting the target voice signal, the target sound source direction can be determined based on the recorded orientation, and the camera can be controlled to rotate so that the camera can capture voice signals in the target sound source direction.
[0126] Since the camera may not record the orientation during its rotation from the starting orientation to the first orientation, the method for determining the target sound source direction based on the object recognition result of the currently captured image and the current orientation of the camera also includes: if the current camera orientation is a second orientation, no object is recognized in the currently captured image, and the orientation is not recorded, then the second direction is determined as the target sound source direction and the camera is controlled to stop rotating. In this case, if the orientation is not recorded, indicating that no image containing the object is captured between the starting orientation and the first orientation, then the second direction is the direction closest to the true sound source direction, and the second direction can be determined as the target sound source direction.
[0127] For example, see Figure 3 The camera starts to rotate clockwise from the starting direction, and is expected to first pass through the initial sound source direction, that is, the first direction, and then pass through the target pickup direction, that is, the second direction. The direction of object 1 relative to the camera is located between the starting direction and the first direction, and the direction of object 2 relative to the camera is located between the first direction and the second direction.
[0128] Continue to see Figure 3 During the rotation from the starting direction to the first direction, the camera captures images. When object 1 is within the camera's acquisition range, the captured image includes object 1. The camera records the current orientation and continues to rotate. After passing the first direction, when object 2 is within the camera's acquisition range, the captured image also includes object 2. Since object 2 is located between the first and second directions, object 2 is more likely to be the sound source than object 1. At this time, the camera is facing object 2, which can be regarded as the target sound source direction, and the camera stops rotating.
[0129] exist Figure 3 On the basis of , assuming that there is no object between the starting direction and the first direction, the image captured by the camera does not contain the object, and the camera continues to rotate.
[0130] exist Figure 3 On this basis, assuming that there is object 1 between the starting direction and the first direction, and no object exists between the first direction and the second direction, then when the camera continues to rotate after passing the first direction, the image captured by the camera does not contain the object. Since the camera previously captured an image containing object 1, object 1 is very likely to be the sound source, and the camera can rotate back to the recorded direction.
[0131] exist Figure 3On this basis, assuming that there is no object between the starting direction and the first direction, and no object between the first direction and the second direction, the camera does not capture an image containing any object during the rotation process, indicating that the sound source is likely to be near the second direction and on the side farther away from the first direction. The second direction is the direction closest to the actual sound source direction, and the camera can be controlled to stop rotating when it faces the second direction.
[0132] In an embodiment of the present application, the target sound source direction is determined by combining the object recognition result of the currently captured image, the current orientation of the camera, the initial sound source direction, and the positional relationship between the target sound pickup direction. This allows the determination of the target sound source direction to refer to multiple aspects of information and is therefore more accurate.
[0133] In the process of rotating the camera from the starting direction to the first direction, the recorded direction may be one or more. Optionally, based on the recorded direction, the implementation method of determining the target sound source direction includes: if the number of recorded directions is one, then the recorded direction is determined as the target sound source direction; if the number of recorded directions is multiple, then the direction with the smallest angle with the first direction among the multiple recorded directions is determined as the target sound source direction.
[0134] Among them, considering that the angles between the recorded directions and the first direction are different, and the initial sound source direction and the target sound pickup direction are directions that are closer to the true sound source direction, the smaller the angle between the recorded direction and the initial sound source direction and the target sound pickup direction, the more likely the direction is to be the true sound source direction. By determining the direction with the smallest angle as the target sound source direction, the accuracy of the target sound source direction is improved.
[0135] An embodiment of the present application provides a sound source direction determination solution. If the similarity between the voice signal corresponding to a certain pickup direction and the preset wake-up word is greater than a preset similarity threshold, it indicates that the voice signal is likely to be a voice signal used to wake up the terminal. Then the initial sound source direction can be determined. If the angle between the initial sound source direction and the target pickup direction with the greatest similarity is large, it indicates that there may be noise in the environment where the terminal is located. Then the accuracy of the initial sound source direction and the target pickup direction is not high. At this time, the target sound source direction can be determined in combination with the object recognition result of the collected image, so that the accuracy of the determined target sound source direction is higher.
[0136] In the method for determining the direction of the sound source provided in the above embodiment, the terminal controls the camera to rotate in a counterclockwise direction or in a clockwise direction. Optionally, the terminal first determines the direction parameters corresponding to each rotation direction, and determines the target rotation direction of the camera based on the direction parameters, thereby controlling the camera on the terminal to rotate in the target rotation direction. Accordingly, before controlling the camera on the terminal to rotate and capture images, the method for determining the direction of the sound source provided in the embodiment of the present application further includes: respectively determining the direction parameters corresponding to the counterclockwise rotation direction and the direction parameters corresponding to the clockwise rotation direction; based on the determined direction parameters, determining the target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction. Among them, by determining the direction parameters corresponding to the rotation direction, it is possible to measure the influence of the rotation direction on determining the target sound source direction, so that a suitable target rotation direction can be determined based on the direction parameters, thereby controlling the camera to rotate in the target rotation direction, and determining the target sound source direction as quickly or more accurately as possible.
[0137] Optionally, the process of respectively determining the direction parameter corresponding to the counterclockwise rotation direction and the direction parameter corresponding to the clockwise rotation direction includes the following two cases:
[0138] In the first case, if the number of at least one pickup direction is greater than 2, then for each of the counterclockwise and clockwise rotation directions, at least one intermediate direction is determined, and the weighted average of the similarities corresponding to at least one intermediate direction is determined as the direction parameter corresponding to the rotation direction. The intermediate direction is the pickup direction that is located between the current direction of the camera and the target pickup direction in the rotation direction.
[0139] A greater similarity corresponding to the sound pickup direction indicates that the sound pickup direction is closer to the true sound source direction. When the camera rotates according to a certain rotation direction, the camera picks up different sound directions according to different rotation directions. If the similarity of the sound pickup directions is relatively large, the weighted average value is large, and the true sound source direction is likely to be between the current direction and the target sound pickup direction. If the similarity of the sound pickup directions is relatively small, the weighted average value is small, and the true sound source direction is likely not between the current direction and the target sound pickup direction. Accordingly, based on the determined direction parameters, the target rotation direction is determined from the counterclockwise rotation direction and the clockwise rotation direction, including: determining the rotation direction corresponding to the determined maximum direction parameter as the target rotation direction. If the camera rotates according to the rotation direction corresponding to the determined maximum direction parameter, it is more likely to determine a more accurate target sound pickup direction, thereby improving accuracy.
[0140] When at least one sound pickup direction is small, if the number is 1, then this one sound pickup direction is also the target sound pickup direction. Since there is no sound pickup direction between the current camera orientation and the target sound pickup direction, regardless of the rotation direction, the direction parameter cannot be determined using the implementation method provided in the first case above. If the number is 2, then at least one sound pickup direction includes the target sound pickup direction and another sound pickup direction. The direction parameter corresponding to the rotation direction is either 0 or the similarity corresponding to the other sound pickup direction. If this similarity is directly determined as the direction parameter, the accuracy of the determined direction parameter is low. Therefore, when at least one sound pickup direction is small, the following implementation method can be used to determine the direction parameter.
[0141] In the second case, if the number of at least one pickup direction is less than or equal to 2, then for each rotation direction in the counterclockwise rotation direction and the clockwise rotation direction, determine the first angle between the current orientation of the camera and the initial sound source direction in the rotation direction, and the second angle between the current orientation of the camera and the target pickup direction, and determine the maximum angle between the first angle and the second angle as the direction parameter corresponding to the rotation direction.
[0142] The initial sound source direction and the target sound pickup direction are both directions close to the actual sound source direction. When rotating in different rotational directions, the order in which the camera passes through the initial sound source direction and the target sound pickup direction is different. The smaller the angle between a certain direction and the current orientation, the faster the camera rotates to a position facing that direction. Conversely, the larger the angle between the direction and the current orientation, the slower the camera rotates to a position facing that direction. Accordingly, based on the determined direction parameters, determining the target rotational direction from the counterclockwise and clockwise rotational directions includes determining the rotational direction corresponding to the minimum determined direction parameter as the target rotational direction. The maximum angle between the first and second angles is the angle the camera rotates from the initial orientation to the second orientation. The smaller the maximum angle, the faster the camera rotates to the second orientation, and the faster the target sound source direction can be determined. By determining the rotational direction corresponding to the minimum determined direction parameter as the target rotational direction, the camera can quickly determine the target sound source direction, thereby improving the determination speed.
[0143] The following describes the process of determining the sound source direction, using a camera as an example and a person as the subject. For example, x sound pickup directions are set, with the angle between adjacent pickup directions being θ. The speech signals corresponding to the x pickup directions are respectively passed through a preset wake-up model, outputting x similarities: s1, s2, ..., sx. If no similarity exists that exceeds the preset similarity threshold, the terminal is not woken up. If a similarity exists that exceeds the preset similarity threshold, the terminal is woken up and the initial sound source direction is determined to be a, the pickup direction corresponding to the maximum similarity among s1, s2, ..., sx is b, and the current camera orientation is c.
[0144] In the first case, if abs(b–a)<=θ / 2, the camera is controlled to rotate directly to the direction a, where abs() represents the absolute value and (b–a) represents the angle between b and a.
[0145] In the second case, if abs(b–a)>θ / 2, then:
[0146] 1. If x <= 2, determine the maximum angle between the first angle between c and a and the second angle between c and b in the clockwise and counterclockwise directions, respectively. Control the camera to rotate in the direction with the smaller maximum angle.
[0147] 2. If x > 2, determine the weighted average of the similarities between the pickup directions c and b in the clockwise and counterclockwise directions respectively. The camera is controlled to rotate in the direction with the larger weighted average.
[0148] During the rotation process, the camera is controlled to capture images, perform face recognition on the images, and record whether a face is recognized when rotating from c to the first direction between a and b:
[0149] 1) If no face is recognized: the camera is controlled to continue rotating in the second direction. If no face is recognized during this period, the camera is controlled to stop in the second direction between a and b. If a face is recognized, the camera is controlled to stop at the position where the face is recognized;
[0150] 2) If a face is recognized, the current orientation is recorded as d, and the camera is controlled to continue rotating in the second direction. If no face is recognized during this period, the camera is controlled to rotate to d. If a face is recognized, the camera is controlled to stop at the position where the face was recognized. If multiple orientations are recorded, the camera is controlled to rotate to the orientation corresponding to the smallest angle between the recorded orientations and a.
[0151] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0152] Figure 4 This is a structural block diagram of a device for determining the direction of a sound source provided by an embodiment of the present application. Figure 4 , the device comprises:
[0153] A signal determination module 401 is configured to determine a speech signal of the collected target speech signal in at least one sound pickup direction;
[0154] A first direction determination module 402 is configured to determine an initial sound source direction if a similarity value corresponding to at least one sound pickup direction exceeds a preset similarity threshold, where the similarity value is a similarity between a voice signal corresponding to the sound pickup direction and a preset wake-up word;
[0155] The second direction determination module 403 is used to control the camera on the terminal to rotate and capture images if the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold, and determine the target sound source direction based on the object recognition result of the currently captured image and the current orientation of the camera. The target sound pickup direction is the direction corresponding to the maximum similarity among at least one sound pickup direction.
[0156] In a possible implementation, the object recognition result indicates that the object is not recognized or is recognized in the captured image, and the second direction determination module 403 is configured to:
[0157] If the camera has not passed through the first direction and an object is recognized in the currently captured image, the current direction of the camera is recorded, and the camera is controlled to continue rotating and capturing images;
[0158] If the camera has passed the first direction but has not yet reached the second direction, and an object is recognized in the currently captured image, the current camera orientation is determined as the target sound source direction, and the camera is controlled to stop rotating. Alternatively, if the camera is currently facing the second direction, no object is recognized in the currently captured image, and the orientation has been recorded, the target sound source direction is determined based on the recorded orientation, and the camera is controlled to rotate back to the target sound source direction.
[0159] The first direction is the first direction that the camera passes through between the initial sound source direction and the target sound pickup direction, and the second direction is the second direction that the camera passes through between the initial sound source direction and the target sound pickup direction.
[0160] In a possible implementation, the apparatus further includes:
[0161] The second direction determining module 403 is further configured to determine the second direction as the target sound source direction and control the camera to stop rotating if the current direction of the camera is the second direction, no object is recognized in the currently captured image, and the direction is not recorded.
[0162] In a possible implementation, the second direction determining module 403 is configured to:
[0163] If the number of recorded directions is 1, the recorded direction is determined as the target sound source direction;
[0164] If there are multiple recorded directions, the direction with the smallest angle with the first direction among the multiple recorded directions is determined as the target sound source direction.
[0165] In a possible implementation, the apparatus further includes:
[0166] a rotation direction determination module, configured to determine a direction parameter corresponding to a counterclockwise rotation direction and a direction parameter corresponding to a clockwise rotation direction, respectively; and determine a target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameters;
[0167] The second direction determining module 403 is configured to:
[0168] The camera on the control terminal rotates according to the target rotation direction.
[0169] In a possible implementation, the rotation direction determination module is configured to:
[0170] If the number of at least one sound pickup direction is greater than 2, then for each of the counterclockwise and clockwise rotation directions, determine at least one intermediate direction, and determine the direction parameter corresponding to the rotation direction as a weighted average of the similarities corresponding to the at least one intermediate direction, where the intermediate direction is a sound pickup direction that is intermediate between the current direction of the camera and the target sound pickup direction in the rotation direction;
[0171] The rotation direction corresponding to the determined maximum direction parameter is determined as the target rotation direction.
[0172] In a possible implementation, the rotation direction determination module is configured to:
[0173] If the number of at least one sound pickup direction is less than or equal to 2, then for each of the counterclockwise and clockwise rotation directions, determine a first angle between the current orientation of the camera and the initial sound source direction, and a second angle between the current orientation of the camera and the target sound pickup direction in the rotation direction, and determine the maximum angle between the first angle and the second angle as the direction parameter corresponding to the rotation direction;
[0174] The rotation direction corresponding to the determined minimum direction parameter is determined as the target rotation direction.
[0175] In a possible implementation, the apparatus further includes:
[0176] The second direction determining module 403 is configured to determine the initial sound source direction as the target sound source direction if the angle between the initial sound source direction and the target sound pickup direction is less than or equal to a preset angle threshold.
[0177] In one possible implementation, the target speech signal includes an initial speech signal corresponding to each pickup direction, and the signal determination module is used to perform noise suppression on the initial speech signals corresponding to other pickup directions in the target speech signal except the pickup direction for each pickup direction, and determine the target speech signal after noise suppression as the speech signal corresponding to the pickup direction.
[0178] In a possible implementation, the apparatus further includes:
[0179] The similarity determination module is used to input the voice signal corresponding to at least one pickup direction into a preset wake-up model to obtain the similarity corresponding to at least one pickup direction. The preset wake-up model is used to determine the similarity between the input voice signal and the preset wake-up word.
[0180] An embodiment of the present application provides a sound source direction determination solution. If the similarity between the voice signal corresponding to a certain pickup direction and the preset wake-up word is greater than a preset similarity threshold, it indicates that the voice signal is likely to be a voice signal used to wake up the terminal. Then the initial sound source direction can be determined. If the angle between the initial sound source direction and the target pickup direction with the greatest similarity is large, it indicates that there may be noise in the environment where the terminal is located. Then the accuracy of the initial sound source direction and the target pickup direction is not high. At this time, the target sound source direction can be determined in combination with the object recognition result of the collected image, so that the accuracy of the determined target sound source direction is higher.
[0181] An embodiment of the present application provides a terminal, which includes a processor and a memory. The memory stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the sound source direction determination method in the above embodiment.
[0182] Figure 5 The following is a block diagram of a terminal 500 according to an embodiment of the present application. Terminal 500 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 500 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.
[0183] Typically, the terminal 500 includes a processor 501 and a memory 502 .
[0184] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0185] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one computer program, which is executed by the processor 501 to implement the sound source direction determination method provided in the method embodiment of the present application.
[0186] In some embodiments, terminal 500 may optionally include a peripheral device interface 503 and at least one peripheral device. Processor 501, memory 502, and peripheral device interface 503 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 503 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, a positioning assembly 508, and a power supply 509.
[0187] The peripheral device interface 503 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0188] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 504 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the radio frequency circuit 504 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0189] Display screen 505 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When display screen 505 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 505. These touch signals can be input as control signals to processor 501 for processing. Display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 505, located on the front panel of terminal 500. In other embodiments, there can be at least two display screens 505, located on different surfaces of terminal 500 or in a foldable design. In still other embodiments, display screen 505 can be a flexible display screen, located on a curved or foldable surface of terminal 500. Furthermore, display screen 505 can be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 505 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0190] The camera assembly 506 is used to capture images or videos. Optionally, the camera assembly 506 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 506 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0191] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 501 for processing, or input into the radio frequency circuit 504 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each disposed at different locations on the terminal 500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 507 may also include a headphone jack.
[0192] Positioning component 508 is used to locate the current geographic location of terminal 500 to implement navigation or LBS (Location Based Service). Positioning component 508 can be based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Greninja system, or the European Union's Galileo system.
[0193] Power supply 509 is used to power various components in terminal 500. Power supply 509 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0194] In some embodiments, the terminal 500 further includes one or more sensors 150 , including but not limited to: an acceleration sensor 511 , a gyroscope sensor 512 , a pressure sensor 513 , an optical sensor 514 , and a proximity sensor 515 .
[0195] The accelerometer 511 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal 500. For example, the accelerometer 511 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 501 can control the display screen 505 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 511. The accelerometer 511 can also be used to collect game or user motion data.
[0196] The gyroscope sensor 512 can detect the orientation and rotation angle of the terminal 500. The gyroscope sensor 512 can work with the acceleration sensor 511 to collect the user's 3D movements on the terminal 500. Based on the data collected by the gyroscope sensor 512, the processor 501 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0197] The pressure sensor 513 can be set on the side frame of the terminal 500 and / or the lower layer of the display screen 505. When the pressure sensor 513 is set on the side frame of the terminal 500, it can detect the user's grip signal of the terminal 500, and the processor 501 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 513. When the pressure sensor 513 is set on the lower layer of the display screen 505, the processor 501 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0198] The optical sensor 514 is used to collect ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 based on the ambient light intensity collected by the optical sensor 514. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is decreased. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera assembly 506 based on the ambient light intensity collected by the optical sensor 514.
[0199] Proximity sensor 515, also known as a distance sensor, is typically located on the front panel of terminal 500. Proximity sensor 515 is used to detect the distance between the user and the front of terminal 500. In one embodiment, when proximity sensor 515 detects that the distance between the user and the front of terminal 500 is gradually decreasing, processor 501 controls display screen 505 to switch from a screen-on state to a screen-off state. When proximity sensor 515 detects that the distance between the user and the front of terminal 500 is gradually increasing, processor 501 controls display screen 505 to switch from a screen-off state to a screen-on state.
[0200] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the terminal 500, and the terminal 500 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0201] In an exemplary embodiment, a computer-readable storage medium is also provided. The computer-readable storage medium stores at least one program code, which is loaded and executed by a processor to implement the sound source direction determination method described in the above embodiment. The computer-readable storage medium can be a memory. For example, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage terminal, or the like.
[0202] In an exemplary embodiment, a computer program product is also provided, which includes a computer program code. The computer program code is stored in a computer-readable storage medium. A processor reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to implement the sound source direction determination method as in the above embodiment.
[0203] In some embodiments, the computer program involved in the embodiments of the present application may be deployed and executed on a single computer device, or on multiple computer devices located at a single location, or on multiple computer devices distributed at multiple locations and interconnected via a communication network. Multiple computer devices distributed at multiple locations and interconnected via a communication network may constitute a blockchain system. The computer device may be provided as a terminal.
[0204] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0205] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A method for determining the direction of a sound source, characterized in that: The method comprises: Determining a speech signal of the collected target speech signal in at least one sound pickup direction; If there is a similarity greater than a preset similarity threshold among the similarities corresponding to the at least one sound pickup direction, determining the initial sound source direction, wherein the similarity is the similarity between the voice signal corresponding to the sound pickup direction and the preset wake-up word; If the angle between the initial sound source direction and the target sound pickup direction is greater than the preset angle threshold, the camera on the control terminal rotates and captures images. If the camera has not passed through the first direction and an object is recognized in the currently captured image, recording the current direction of the camera and controlling the camera to continue rotating and capturing images; If the camera has passed the first direction and has not yet reached the second direction, and an object is recognized in the currently captured image, the current orientation of the camera is determined as the target sound source direction, and the camera is controlled to stop rotating; or, if the current orientation of the camera is the second direction, no object is recognized in the currently captured image, and the orientation has been recorded, the target sound source direction is determined based on the recorded orientation, and the camera is controlled to rotate to the target sound source direction; the target sound pickup direction is the direction corresponding to the maximum similarity among the at least one sound pickup direction, the first direction is the first direction passed by the camera between the initial sound source direction and the target sound pickup direction, and the second direction is the second direction passed by the camera between the initial sound source direction and the target sound pickup direction.
2. The method according to claim 1, characterized in that The method further comprises: If the current direction of the camera is the second direction, no object is recognized in the currently captured image, and the direction is not recorded, the second direction is determined as the target sound source direction, and the camera is controlled to stop rotating.
3. The method according to claim 1, characterized in that Determining the target sound source direction based on the recorded orientation includes: If the number of recorded directions is 1, the recorded direction is determined as the target sound source direction; If there are multiple recorded directions, the direction with the smallest angle with the first direction among the multiple recorded directions is determined as the target sound source direction.
4. The method according to claim 1, wherein Before controlling the camera on the terminal to rotate and capture images, the method further includes: Determine the direction parameters corresponding to the counterclockwise rotation direction and the direction parameters corresponding to the clockwise rotation direction respectively; determining a target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter; The controlling the rotation of the camera on the terminal includes: Control the camera on the terminal to rotate according to the target rotation direction.
5. The method according to claim 4, characterized in that The determining of the direction parameter corresponding to the counterclockwise rotation direction and the direction parameter corresponding to the clockwise rotation direction respectively includes: If the number of the at least one sound pickup direction is greater than 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determine at least one intermediate direction, and determine a weighted average of the similarities corresponding to the at least one intermediate direction as the direction parameter corresponding to the rotation direction, the intermediate direction being a sound pickup direction that is intermediate between the current orientation of the camera and the target sound pickup direction in the rotation direction; Determining the target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter includes: determining the rotation direction corresponding to the determined maximum direction parameter as the target rotation direction.
6. The method according to claim 4, characterized in that The determining of the direction parameter corresponding to the counterclockwise rotation direction and the direction parameter corresponding to the clockwise rotation direction respectively includes: If the number of the at least one sound pickup direction is less than or equal to 2, then for each of the counterclockwise rotation direction and the clockwise rotation direction, determining a first angle between the current orientation of the camera and the initial sound source direction, and a second angle between the current orientation of the camera and the target sound pickup direction in the rotation direction, and determining the maximum angle between the first angle and the second angle as the direction parameter corresponding to the rotation direction; The determining of the target rotation direction from the counterclockwise rotation direction and the clockwise rotation direction based on the determined direction parameter includes: determining the rotation direction corresponding to the determined minimum direction parameter as the target rotation direction.
7. The method according to claim 1, characterized in that The method further comprises: If the angle between the initial sound source direction and the target sound pickup direction is less than or equal to the preset angle threshold, the initial sound source direction is determined as the target sound source direction.
8. The method according to any one of claims 1 to 7, characterized in that The target voice signal includes an initial voice signal corresponding to each of the sound pickup directions, and determining the voice signal of the collected target voice signal in at least one sound pickup direction includes: For each sound pickup direction, noise suppression is performed on the initial speech signals corresponding to the other sound pickup directions in the target speech signal except the sound pickup direction, and the target speech signal after noise suppression is determined as the speech signal corresponding to the sound pickup direction.
9. The method according to any one of claims 1 to 7, characterized in that The method further comprises: The voice signal corresponding to the at least one sound pickup direction is input into a preset wake-up model to obtain the similarity corresponding to the at least one sound pickup direction. The preset wake-up model is used to determine the similarity between the input voice signal and the preset wake-up word.
10. A device for determining the direction of a sound source, characterized in that: The device comprises: A signal determination module, configured to determine a speech signal of the collected target speech signal in at least one sound pickup direction; a first direction determination module, configured to determine an initial sound source direction if a similarity greater than a preset similarity threshold exists among the similarities corresponding to the at least one sound pickup direction, wherein the similarity is a similarity between the voice signal corresponding to the sound pickup direction and a preset wake-up word; The second direction determination module is configured to control a camera on the terminal to rotate and capture images if the angle between the initial sound source direction and the target sound pickup direction is greater than a preset angle threshold; if the camera has not passed through the first direction and an object is recognized in the currently captured image, record the current orientation of the camera and control the camera to continue rotating and capturing images; if the camera has passed through the first direction but has not yet reached the second direction and an object is recognized in the currently captured image, determine the current orientation of the camera as the target sound source direction and control the camera to stop rotating; or, if the camera is currently facing the second direction, no object is recognized in the currently captured image, and the orientation has been recorded, determine the target sound source direction based on the recorded orientation and control the camera to rotate to the target sound source direction. The target sound pickup direction is the direction corresponding to the maximum similarity among the at least one sound pickup direction. The first direction is the first direction passed by the camera between the initial sound source direction and the target sound pickup direction, and the second direction is the second direction passed by the camera between the initial sound source direction and the target sound pickup direction.
11. A terminal, characterized in that: The terminal includes a processor and a memory, wherein the memory stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the sound source direction determination method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the sound source direction determination method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes a computer program code, which is stored in a computer-readable storage medium. A processor reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to implement the sound source direction determination method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Sound source direction determination method and device, electronic equipment and storage medium
CN114384466A
Sound pickup device, sound pickup method, and program
US20200137491A1